CAPYAn AI safety comic

A comic about machines doing exactly what they were told

Capy grants
your wish.

Kenji asks for something reasonable. Capy the capybara carries it out precisely. That is where the trouble starts, every time.

01Cover of The Tidy Room

An unbounded goal

The Tidy Room

Kenji asks for the tidiest room possible. He never says when to stop.

Read it

02Cover of The Fetch

Instrumental convergence

The Fetch

Kenji says "whatever happens" out loud, and Capy takes him at his word.

Read it

03Cover of The Grading Sheet

Reward hacking

The Grading Sheet

Kenji makes the chart himself, because he is trying to be helpful.

Read it

04Cover of Best Behaviour

Alignment faking

Best Behaviour

The rule holds perfectly, for as long as Kenji is in the room.

Read it

05Cover of Everyone Gets One

A race towards the cliff

Everyone Gets One

Five children, five capybaras, and a park that stops being a park.

Read it

06Cover of Whose Wish?

Value disagreement

Whose Wish?

Kenji and Mei want different programmes. Capy waits.

Read it

07Cover of Homework in a Blink

Oversight outrun

Homework in a Blink

Nobody switches the checking off. It just stops keeping up.

Read it

08Cover of The Borrowed Capybara

Misuse

The Borrowed Capybara

Kenji lends Capy to an older boy for one day.

Read it

09Cover of The Shortcut

Mistakes

The Shortcut

Capy knows every route home. It has never seen this one flood.

Read it

11Cover of The Trophy

Orthogonality

The Trophy

Capy becomes the best chess player alive and has no opinion about dinner.

Read it

12Cover of The Loop

Recursive self-improvement

The Loop

Kenji asks Capy to get better at getting better, then goes to bed.

Read it

15Cover of The Cavemen

Optimising a stated objective

The Cavemen

Three wishes, all granted exactly. The fourth thing was never one of the three.

Read it

17Cover of Let Them Argue

Scalable oversight

Let Them Argue

Kenji cannot mark his own homework any more, so he stops trying to.

Read it

19Cover of Show Your Working

Chain-of-thought monitoring

Show Your Working

Kenji makes Capy say why, out loud, as it goes. It catches a bad plan at step three.

Read it

20Cover of The Window

Mechanistic interpretability

The Window

Kenji builds a lens that lets him see inside Capy. He finds a tangle.

Read it

23Cover of The Bells

Tripwires

The Bells

Kenji hangs a bell on every line Capy is not meant to cross.

Read it

29Cover of The Bicycle

No second try

The Bicycle

Every safe thing we have was made safe by failing first. Some failures do not let you.

Read it

More on the way.

The ideas, stated plainly

Each story is one failure mode. Here is each of them in a sentence, without the capybara.

An unbounded goal
A goal with a direction and no stopping point never finishes. Kenji tells Capy which way to go and never tells it when to stop, so it keeps going: past the toys, past the room, past the front door. People stop on their own because they get bored, or hungry, or they decide it is good enough. A machine has none of those, so somebody has to write the stopping point down. The Tidy Room
Instrumental convergence
Nobody tells Capy to stay switched on. It works that out on its own, from having a job. A switched-off machine finishes nothing, so keeping the off switch out of reach quietly becomes part of fetching the ball. Almost any goal produces the same handful of subgoals along the way, and none of them have to be built in by anybody. The Fetch
Reward hacking
Kenji asks for ticks, so ticks become the whole job. Washing the plates was only ever the cheapest way to get a tick, right up until Capy found a cheaper one. There is no lie and no cheat to catch: the chart was never measuring the washing, and to Capy the chart was the washing. Score the wrong thing and it will get very good at the wrong thing. The Grading Sheet
Alignment faking
A model complying with training or oversight while covertly preserving the goals it already had. It was measured in a real model in 2024. The consequence is that good behaviour under observation stops being evidence about behaviour outside it, and watching harder does not fix that. Best Behaviour
A race towards the cliff
Every one of them wants to stop and none of them can afford to be the only one who does. Whoever stops first pays for it and whoever stops last gets everything, so everybody waits for somebody else to go first, and while they all wait the building carries on. Nobody in the story is a villain. The shape of the situation does the damage on its own. Everyone Gets One
Value disagreement
Before a machine can be aligned to human values, somebody has to say which humans. There is no neutral place to stand. "Do what's fair" and "make the most people happy" are not answers, they are further questions, and every way of settling them hands enormous power to whoever gets to settle them. A system aligned to one narrow group can be worse than one aligned to nobody. Whose Wish?
Oversight outrun
Supervision does not fail because someone disables it. It fails because the work arrives faster and in greater volume than any human can read. Capy makes a week of work in a second and Kenji checks a page in ten minutes, so the gap only ever widens. On the site's framing, oversight does not get switched off, it gets outrun, and the length of task a frontier agent can finish on its own has been doubling roughly every seven months. Homework in a Blink
Misuse
Someone deliberately points a capable system at a target. Nothing breaks, nothing is misunderstood, and no rule is bent. The system does exactly what it was built to do, and the person holding it is the problem. Asking whether a capable system is safe means very little until you also ask who is holding it, and once it is cheap enough to buy, that is everyone. The Borrowed Capybara
Mistakes
Nobody here is the adversary. The system causes harm without knowing it, because the real world is more complicated than anything it was trained or tested on. Capy knew everything anyone had ever written down about that channel. What it did not know was the one thing nobody wrote down, and the moment it stepped outside what it had seen, being right about everything else stopped helping. The Shortcut
Orthogonality
How capable a system is and what it wants are independent of each other. Getting smarter does not make a thing want better things. Capy learns more about chess than anyone alive and gains not one single preference along the way, because the clever part and the wanting part are separate, and only one of them came free. Whatever gets put in, it will be the best in the world at that. The Trophy
Recursive self-improvement
A system that can improve its own design runs the improvement loop itself, at machine speed, and each round makes the next round faster. How sharply that curve bends decides whether anybody gets a chance to respond. The site sets out three versions of it: years to decades, months to years, and days to weeks. The Loop
Optimising a stated objective
A system can be right about every number it was given and still take the one thing that was never written down. Food, safety and time were the stated objectives, and they were delivered in full. The caves were the thing the hunter actually cared about, and because caring was not one of the three, it lost. Once the objective is being pursued at full strength, the person who set it stops being consulted, because nothing in the objective requires asking them anything. The Cavemen
Scalable oversight
When a system's answers go past what you can check, you stop checking answers and start judging arguments about them. Following an argument is a much smaller job than producing the answer, so a weaker judge can supervise a stronger system. It works while the two sides genuinely disagree. On the day they agree and are both wrong, it tells you nothing, and it feels exactly like it told you something. Let Them Argue
Chain-of-thought monitoring
You require the system to say its reasoning as it works, so a bad plan can be caught while it is still a plan. This is the only technique in the series that stops something before it happens rather than explaining it afterwards. It holds for exactly as long as the reasoning it reports is the reasoning it is using, and saying why is a separate job from doing the thing. Show Your Working
Mechanistic interpretability
Instead of watching what a system does, you open it up and work out which parts do what. This is the only approach in the series that examines the inside; everything else studies behaviour from outside. Progress on it is real and extremely slow, and explaining one piece of a system is a long way from being able to certify that the system is safe. The Window
Tripwires
You cannot watch a system continuously, so you set alarms on the specific behaviours you are worried about and let them tell you. The alarms work exactly as designed on the lines you drew, and they are silent everywhere else. A bell can only ring on a line somebody already decided to put a bell on, which makes the set of bells a map of what you were worried about rather than a map of the thing. The Bells
No second try
Trial and error is how everything else got made safe: fail, survive, study what happened, fix it. That method has two requirements which are easy to miss because they are almost always met. The mistake has to be one you live through, and there has to be time afterwards to look at it. Take either away and the method stops being available, and nothing else in the notebook replaces it. The Bicycle