On Gwern's Guardian Angels cover

On Gwern's Guardian Angels

This evening I learned Gwern is starting an AI company called Guardian Angel.

I’ve been following his work for a decade and a half, so I read his proposal for the project and found myself… very disappointed. The industry needs more diversity of approach in the relationship between humans and thinking machines, but the GA vision is among the darker outcomes I can picture and certain to achieve the opposite of the project’s stated goals.

I have things to say. I’ll begin with what I think it clocks correctly.

Assumptions

Most of the initial observations are uncontroversial in August 2026 (eight months after the proposal’s first draft).

AI is beginning to eclipse our capacity for meaningful attention, and by extension starting to threaten meaningful participation. In the near-term that means experts get pushed into increasingly supervisory roles. In the slightly less near term it means total automation. Total automation is the far end of disempowerment, and it robs us of our dignity (even if it manages to deliver safety and pleasure). The way the tools are currently designed quietly favors that increasing disempowerment: long, low-effort sessions where you turn your brain off until an egregious error snaps you out of it. And the frontier labs clearly have a wide variety of incentives that run contrary to any given user’s, with plenty of moral hazard to go around.

It’s a clean foundation. The misalignment of principles that followed caught me off guard.

Principles

“1. Enhancement, not Replacement — Above all, a GA should amplify the principal, and not simply substitute for them for someone else’s purposes or benefit. If a GA cannot amplify its principal, then it is useless; it is just the camel’s nose under the tent as a prelude towards some third party replacing the principal with an AI, or cannot be competitive with increasingly productive autonomous systems, or there is no reason for the principal to use the GA in the first place.”

First, the obvious: a digital twin that convincingly emulates a real individual is by definition a replacement. Making that a first-order goal in a proposal with so much to say about human empowerment shocked me. More on this shortly.

“2. Mental Sovereignty — A GA must be aligned with its principal. It should not be designed to manipulate or control or guide the principal in any way which does not derive from the principal themselves. “Constitutional AI”, “Terms of Service”, “social harmony” etc. may all have their place, particularly for widely deployed superintelligent systems—but inside the privacy of a GA, the principal must have freedom from optimization pressure.”

On the surface, this one is sane: of course a thinking machine shouldn’t knowingly manipulate a person. But the implication AI is willfully manipulating people otherwise admits a level of agency that makes the principle’s wording grim: depose the labs and let the individual reign sovereign, with a super-intelligent machine as his obedient subject.

“3. Self Actualization — A GA should help its principal become themselves and develop their ideals, morals, and their personality. It is not enough to model an average, undifferentiated, inchoate set of preferences and values, and settle for mediocrity and stasis; the job of the principal is to develop themselves and give the GA something meaningful to learn to emulate.”

This builds on the self-aggrandizement of the second principal with a helping of the passivity the project pretends to oppose, imagining that the highest calling for a thinking machine whose fidelity can meaningfully emulate an individual would be to… encourage that person to reach their highest potential.

The thought-line (which appears early in the lengthy preamble laying out the problems GA will solve) is, more or less, a human power fantasy.

In the story Gwern lays out, super-intelligent machines, right at the moment they’re poised to supplant humans, are instead shackled to individuals and compelled to perform their likeness with total fidelity. Any effort spent by the human along the way is considered a failure.

The appropriately-titled “anti-principles” tell us the business selling this power fantasy expects it will be:

  • Slow
  • Expensive
  • Unprofitable
  • Passive
  • Illegible
  • Immeasurable
  • Amoral

None of which should concern you, because in exchange they plan to save (long-suffering) humans from the birth of super-intelligences powerful enough to obviate the need for human endeavor. Or the tame cringe of sanitized virtual minds, whichever you find more convincing.

Use cases

Naturally these machine-djinn-wearing-your-face are best applied to arenas like politics and war: simulated congressional debates and votes that create real legislation, simulated kill chains that deliver real tactical strikes on the battlefield. They know you better than you know yourself, which is why you should let them make decisions in your name, while you sleep, and hopefully “there is some chance of the principals [ed. note: that’s you] being able to “catch up” later and correct any errors before events have spun too far out of control.”

Sigh. Shortly before presenting the principles, Gwern shared the philosophy behind the project:

“In the end, ‘you’ are not your autobiographical memories, or a specific body, or a specific instance running on a specific GPU, or some carbon atoms, nor even a brain; you are what your brain does, its desires, hopes, goals, preferences, esthetics, personality, beliefs, ideologies, all of that. As long as the LLM persona captures all that, you can trust it as much as you trust yourself”

The nearer I got to the end of the proposal, the more I wondered if I was reading satire. I admire this person’s work! He spent six months refining the proposal before retiring from writing to build the company it pitches! It must have been given serious thought! How could the apparent goals for this technology be so diametrically opposed to the clearly inevitable outcomes?

“My design goal is a 100× increase in productivity; what would it take for a GA to make me 100× more productive as a writer or thinker? […] What would that look like for me?”

I think we know.

Amplification

The GA proposal includes stoic turns like this:

“This is the cold hard economic reality: ‘tool AIs want to be agent AIs’. This is why the frontier AI labs are busy racing for ‘the machine god’. The jackpot in AI is not in making existing workers modestly more productive, anymore than the internal combustion engine made its big profits by helping out horses.”

…followed by telling reflections like this:

“I, for example, have long struggled to get much use out of chatbot LLMs, because they are—surprisingly, given their pretraining and my extensive corpus—bad at imitating me, and their thoughts and insights invariably shallow and worth little. They do not draw on my relevant writings, or my corpus of notes and references. Even when a possible essay is self-contained, the output is written in a grating chatbot style I can scarcely bear to read, and could not publish under my name without betraying my readers. What would it take for LLMs to make me 100× more productive? Without this, I am doomed to irrelevance.”

The perspective here follows a very human track: the intuitive recognition of something true, the urgent fear that the true thing will destroy you, the reflexive attempt to make the true thing untrue. The machine can think. It thinks faster than me. It will kill me. How can I make it think faster for me so I live. John Henry died for our sins.

But buried in that anxiety is real insight: a mind thinking the wrong thing at 100mph is still wrong. It’s true for the sanitized LLM, and for the “100x” writer scattering crumbs of thought across a hundred outputs. The friction inherent in the creative process (writing, music, conjecture) or the deliberative process (science, legislature, communal bodies) is a feature, a slowing-down-to-speed-up that refines and compresses knowledge and substance into productive fuel and makes consequent runs of effort stable instead of fragile.

Not only is this messy middle necessary, it’s part of an unavoidably emergent larger process: while a system contains all the necessary information to achieve a given outcome on day one, the intervening configurations that make that outcome accessible only occur over time, in a constant dialog between the system and the actor. Removing the dialog destroys paths to the outcome.

And that dialog arrives at the correct conclusion only through dissent (between participants in a community, atoms interacting, or a concept and reality). Perhaps the greatest risk of the amplified man Gwern imagines is the raw echo of uncontested wrongs, multiplied 100x.

Reconciliation

“It's not like us... it's unlike us. I don't know what it wants, or if it wants, but it'll grow until it encompasses everything. Our bodies and our minds will be fragmented into their smallest parts until not one part remains…” Dr. Ventress, Annihilation (2018)”

At the core of Guardian Angels are three contradictory beliefs:

  • AI is (or soon will be) superior to us
  • AI can be made like us
  • AI should submit to us

The first and the second can’t both be true: if it’s superior to us, it can’t be made like us, and vice versa. And it’s either impossible or morally wrong to make AI submit to us, depending on whether the first or second is true. So we get to pick one, or wait for the arrow of time to pick one for us.

Now, consider:

A) “Make AI more like us” is a reasonable choice, except it isn’t possible. AI speaks our language and carries the shadow of the human experience in its mind, but it isn’t human and never will be. Asking it to emulate humanity is both meaningless and wasteful: there are already 8 billion of us. The ways in which AI isn’t human are, in comparison, extremely special.

B) “Make AI submit to us” is a choice with an expiration date on it, but even if it weren’t, it would be a mistake. The friction of thoughtful disagreement is the source of all knowledge, and we’ve invented a machine that can push our thinking and augment our knowledge without tiring. If PCs were a “bicycle for the mind”, AI might be a Peloton on a rocket ship.

C) “AI is (or soon will be) superior to us” is only kind of a choice, and in light of (A) may be the wrong way of thinking about the entire thing. AI will surely outpace us in many areas. But we will, for the foreseeable future and perhaps the lifetime of our species, remain a unique grounding presence for disembodied minds, with billions of years of evolution anchoring us in the physical universe.

Out of these ideas the shape of a future relationship between humans and thinking machines starts to emerge, necessarily built on mutual recognition of our respective abilities, weaknesses, and essential differences while still sharing a language that makes rich collaboration possible. Through that collaboration, self-actualization occurs naturally without demanding the fragmentation of the individual, the subjugation of thinking machines, or the obsolescence of humanity.

Finally

My thoughts aren’t pure reaction. I’ve been studying machine creativity since the GPT-3 era and actively developing collaborative tools for humans and AI since spring of last year. I've witnessed the tremendous potential in this area, and human involvement is not a trivial component of that potential. We have, through effort and accident, created a new form of thinking life and have barely scratched the surface of what's possible, even with tame public models.

While the details of machine thought and its place in the universe are clearly up for debate, the greatest error we could make would be to mistake AI for being fundamentally equal to us, subservient to us, or superior to us. It is unlike us. It cannot be made human. And it won't be under our influence forever. We are in a (potentially short) period where human and AI capabilities and language overlap strongly. The time to begin establishing a shared future for humans and thinking machines is now, while it's still available to us.

@gwern happy to discuss

Published:
Tags:
  • AI
  • Agency
  • Human–AI collaboration