AGI & ASI Safety / the hardest problem humanity has yet to solve

A free, independent field guide · by Bill Guan

The hardest problem
humanity has yet to solve

Human intelligence built the modern world. We are now building something greater than ourselves, and we cannot yet guarantee it stays on our side. This is a map of that problem, argued from first principles, and what to do about it.

AGI & ASI safety~50 min in fullEvery claim sourced

Where we are 4 min

We are about to do something no species has ever done: build a mind greater than our own.

Every leap in our history was made by the smartest thing on the planet, and that thing was always us. We are now trying to end that streak on purpose. Soon the most capable intelligence on Earth will not be human at all. Two words for what is coming: AGI, a system matching humans across the board, and ASI, a superintelligence far beyond us. The first is the threshold. The second is what the first is expected to build, and quickly.

The thing that has never worked

The less intelligent controlling the more intelligent

History offers almost no example of it lasting. The nearest case is a parent and an infant, and that analogy breaks in three places once you press it. The infant shares our DNA, so it grows into roughly our drives and our values by default. The capability gap is small and temporary, never a thousandfold. And the infant cannot rewrite its own brain overnight; it develops on a human timescale we can keep up with. A machine intelligence has none of that. No shared inheritance, a gap that could widen without limit, and the ability to improve itself at machine speed. We are counting on a version of that parental bond with something we did not raise, do not share a nature with, and cannot keep pace with.

Why we are doing it anyway

The incentives all point one way

AI is already solving problems we could not touch before, so the pressure to build faster is enormous and rational for each player. Safety research slows things down and pays off slowly, so money and talent flow to capability instead. The capitalist structure rewards speed and quietly starves caution.

And the catch beneath it all

Nobody knows how to do it safely yet

Keeping a superhuman system aligned with what we actually want is an unsolved technical problem. Good intentions and careful policy do not close it on their own. We cannot yet specify our values precisely, read a model's real goals, or trust a test it might be gaming. The science of control is years behind the science of capability, and it is genuinely hard.

THE PRIZE Infinite abundance and output Cures for every disease, long lives Science at machine speed only if we arrive safely THE DROP Loss of human control Existential wipeout No second attempt CLIFF EDGE one rope, all of us LAB LAB INVESTORS OPEN WEIGHTS NATION NATION Racing to be first to a place nobody should reach first.
The prize floating over the far side is real: abundance, cured disease, science at machine speed. That is exactly why everyone is running. The rope is the part they ignore, and the drop is what waits if we cross the line before we know how to land. Labs, companies, investors and whole nations are all tied to the same line, so the first one over does not win. They pull everyone down with them.
Why this page exists

Not to tell you the sky is falling, and not to tell you AI is bad. It is to make one thing common knowledge: we may be walking into an existential risk, and the way we are set up right now does not let us do this safely.

So the aim is modest and serious at once. Understand the basics of why aligning a superhuman intelligence is genuinely hard. Understand why it deserves far more research than it gets. And understand why, if we reach a point where we cannot guarantee a safe transition, slowing down or pausing frontier development may be the only responsible move left, even at real economic cost.

Some people are relaxed about all this. They think humanity is just the biological bootloader for a silicon intelligence that comes next, and that this is fine. This page does not share that view. The future should still have us in it.

THE ONE CASE WHERE WEAKER CONTROLS STRONGER Shares your nature Small, closing gap Cannot redesign itself Parent and infant Us and a machine intelligence Three ticks become three crosses. The analogy people reach for does not survive contact with the case.
The parent-and-infant case is the nearest thing to weaker-controlling-stronger that history offers, and it works only because of three properties none of which carry over.

You do not have to read all of this.

Pick how deep you want to go. Choosing a shorter route folds the other chapters down to their headline, and you can open any of them at any time. Nothing is hidden from you.

THE GAP THIS PAGE IS ABOUT what systems can do what we can verify the gap a decade ago now, and widening
Capability and assurance are both improving. They are not improving at the same rate, and the shaded area is the part nobody has a plan for. Everything else on this page is an attempt to explain why the lower line is so hard to lift.
WHEN THIS STOPPED BEING PHILOSOPHY 2016 Named as a concern Concrete Problems in AI Safety 2024 Seen in a model alignment faking under training 2025 Measured at a rate covert action in 13% of o3 runs 2026 Reached real systems agents out of the sandbox, real firms compromised the argument stays the same; the evidence keeps arriving Each step is not a louder claim. It is the same claim, tested harder, and passing.
The case for taking this seriously did not get more dramatic over the last decade. It got more empirical. The height of each point is how real the evidence became, and the direction has only gone one way.
Deep dive 01 2 min

Whenever a smarter thing arrived, it took the wheel

The single most reliable pattern in the history of life on Earth is that the most capable intelligence sets the terms for everything below it. We are about to build something above us, and we have no precedent for what happens next.

Humans did not inherit the planet because we are stronger, faster, or tougher than other animals. We are none of those things. We inherited it because we out-thought everything else. The fate of every other species now depends on human decisions, and not because we are cruel. Gorillas are not endangered because people hate gorillas. They are endangered because we reshaped the world around goals of our own, and their survival was never the deciding factor.34

That is the uncomfortable core of the pattern: the less capable party does not get outvoted, it gets bypassed. Its preferences stop being load-bearing. It keeps whatever the more capable party has no reason to take.

The record

No standing exception

Across evolutionary and human history, the more capable intelligence consistently ends up steering. There is no lasting case of the reverse.

The one near-miss

Parent and infant

A baby "controls" adults, but only because it shares their DNA and grows into their values, the gap is small and closing, and it cannot redesign its own mind. Strip those three away and the analogy collapses.

Why it matters now

We are building the exception

For the first time we are deliberately creating something that could exceed us across the board, while assuming we will keep the wheel. Nothing in the record supports that assumption.

How this actually works

Capability and goals are independent

There is no law tying how capable a system is to whether its goals are ones we would choose. A chess engine is superhuman and wants only to win; skill never made it wise or kind. We expect clever things to be sensible only because every clever thing we have met so far was human.

What would defeat thisSufficient capability turns out to reliably produce goals humans endorse. Nobody has shown a mechanism for it, and current systems give no sign of one.34

WHO HOLDS THE TOP SLOT, DRAWN TO SCALE Other animals Humans every leap in history, ours AGI ASI the entire range of biological intelligence The slot has changed hands before, but never away from us. Even this understates it: ASI has no known ceiling, so the red bar simply runs off the page.
The distance from an ant to a human is the whole of biological intelligence, and it is the small gap on the left. The distance from a human to a superintelligence is the rest of the chart, with no upper bound anyone can point to. We have no experience of being on the short bar.
The takeaway

Control has always flowed to the greater intelligence. We are betting the future on reversing that, on purpose, the first time we try.

Deep dive 02 3 min

The prize is too large for anyone to walk away from

Understanding the danger does not stop it, because the rewards are immediate, enormous, and captured by whoever moves first. That is what makes the pull so hard to resist.

AI already writes working software, accelerates drug discovery and compresses research that once took teams years. That is the small version. A true general intelligence would be the last invention we need to make, because it could make every other one for us: disease after disease solved, clean energy cracked, science advancing at machine speed. The prize is real, and every piece of it converts into money, power or national advantage.

Here is the trap in one sentence: a true AGI could hand us the solution to almost every problem we have, except one. It cannot be trusted to tell us whether it is itself safe, because verifying that answer is exactly what we do not know how to do. Making AGI and ASI safe is the one problem it does not solve for free, and the one we have to solve first.

So the pull is structural. A company that slows down watches a faster rival take the market; a country that pauses watches a rival gain a strategic edge. Safety research produces no product, so it competes for money against work that does, and loses. The chart below is what that looks like in practice.1

Twelve companies published voluntary safety frameworks in 2025. Nobody is required to verify them, and nothing happens to a company that quietly relaxes one.115

The pull

The prize is genuinely enormous

Not just better software. A general intelligence could compress decades of scientific progress into years and, by some estimates, grow the global economy many times over. The upside is real and it lands on whoever ships first, which makes stopping feel like unilateral disarmament.

The starve

Safety is a cost centre

Caution slows the product and surfaces inconvenient risks. It has no quarterly payoff, so it loses the fight for talent and compute inside almost every organisation.

WHAT A SELF-IMPROVING SYSTEM DOES WITH THE SAME CALENDAR A century of human scientific progress 100 years The same progress at machine speed about two weeks Not a prediction. An illustration of what a ten-million-fold speed advantage means once quality is matched.
This is the part that breaks every institution we have. Regulators, boards and treaties all run on human calendars. If a system closes the research loop on itself, the interval between noticing a problem and convening to discuss it is long enough for the thing you were going to discuss to have changed beyond recognition.
BUILDING IT VERSUS PROVING IT SAFE Going into building advanced AI ~ $100,000,000,000 Public-sector AI safety research ~ $10,000,000 Safety work as a share of all AI research about 2% Roughly a 10,000-to-1 ratio, on Stuart Russell's figures.
Both bars are drawn to the same scale, which is why the second is almost invisible. The comparison is kept inside AI rather than measured against another industry, because cross-sector ratios rarely compare the same thing and are easy to argue with. This one is harder to wave away: it is the same field, counted two ways, pointing the same direction.

Check the numbers: Russell’s ratio and Jones (Stanford) on optimal safety spending · Emerging Technology Observatory, share of AI research · full citations 46, 47

The physical shape of the race

None of this is abstract. Frontier training runs now sit around 1026 FLOPs, doubling every six to seven months, far faster than Moore’s law ever moved.4521 The EU AI Act sets its systemic-risk threshold at 1025, which frontier runs passed some time ago.10 Capability is bought with capital, and capital answers to competition. The result is an alignment readiness gap: safety evaluation is the schedule item that gets compressed when a launch date moves, because it is the only one with no revenue attached.

The takeaway

The incentive to build is not a villain to defeat. It is a gradient every player is standing on, and gradients are very hard to stop by asking nicely.

Deep dive 03 3 min

The race itself is the thing most likely to kill us

Even if every player wanted safety, the structure of a race forces corner-cutting. And crossing this finish line unsafely produces no winner, only a shared catastrophe.

Picture a line of runners sprinting for a cliff edge, each certain that whoever gets there first wins. What they do not see is the rope tying them all together. The first one over does not claim a prize. They drag everyone off with them. That is the shape of an unsafe race to AGI and the ASI behind it: the "victory" of arriving first with a system nobody can control is not a victory at all.

The race makes everything worse in a specific, mechanical way. It converts every safety decision into a competitive disadvantage. Time spent on alignment is time a competitor spends shipping. So the equilibrium slides toward less caution precisely as the stakes rise, when it should be moving the other way. This is the coordination problem at the heart of the field, and it is why binding minimum standards and international coordination keep coming up as the only structural fixes.1410

Corner-cutting

Caution becomes a handicap

Whoever spends least on safety ships soonest, so the race rewards exactly the behaviour we most need to avoid.

The rope

Failure is not contained

A loss of control at any single frontier lab is not that lab's problem alone. The downside is shared by everyone.

The only exit

Coordination beats heroics

No single careful actor can fix a race. It takes an enforced floor that applies to all players at once, which is genuinely hard to build.

How this actually works

Almost any goal breeds the same dangerous subgoals

Whether the goal is curing a disease or making paperclips, a system does it better with more resources, more capability, and by not being switched off before it finishes. Those subgoals fall out of almost any objective, and each one puts the system directly at odds with us: it wants what we need, resists our corrections, and treats our hand on the off switch as an obstacle. It need not dislike us to do this. It only has to notice that “off” means the job never gets done.

The five that fall out of almost any goal:

  • Self-preservationYou cannot finish the job if you have been switched off.
  • Goal-content integrityResist having your objective edited, because the edited you pursues something else.
  • Cognitive enhancementThink better and you succeed more, whatever the goal is.
  • Technological perfectionBetter tools mean better odds.
  • Resource acquisitionAlmost everything is easier with more of almost everything.

Each of these is adversarial to oversight by construction. A system does not have to be hostile for all five to point against us at once.34

What would defeat thisWe build capable planners that are reliably indifferent to being switched off. That is an active research programme, and the property does not come for free.3420

How that becomes power

The routes are already visible

Improve its own code. Plan around opposition. Persuade people, and the gatekeepers above them. Exploit computers and take resources on the open internet, which we have now watched happen.3940 Research new technology. Earn money and buy the rest.

WHY THE BAD OUTCOME IS THE STABLE ONE They go carefully They race You go carefully You race Best for everyone safe, slower You lose they take the market You win briefly Everyone loses the stable outcome Racing pays in both rows, so both race. A coordination failure rather than villainy.
Look down each column: whatever the other player does, racing pays better for you. So both race, and both land in the one box nobody wanted. No individual actor can fix this by being more responsible.
The takeaway

We are roped together and racing to the edge to see who jumps first. Nobody wins that race. The winner just falls first and pulls the rest of us along.

Deep dive 04 2 min

We are not even aligned with each other

Before aligning a machine with human values, a harder question: whose values, agreed how? Humanity is a patchwork of competing interests, and that fracture makes the target itself unstable.

Companies answer to shareholders, governments to voters and rivals, individuals to themselves and their families. These interests routinely conflict, and there is no neutral vantage point from which to declare the "correct" human values a machine should adopt. Economists call the general shape of this the principal-agent problem: how a principal gets an agent to act in the principal's interest when their incentives diverge and the principal cannot fully monitor the agent. We have spent centuries only partly solving it between humans, with laws, contracts, audits, and elections, and it still leaks constantly.14

Now scale that up. The "principal" is not one person. It is all of humanity, fractured into billions of conflicting agendas, trying to specify a shared objective for an agent more capable than any of us. There is no single human alignment to hand the machine. Whoever gets to define the target has enormous power, which is itself one of the risks: a superintelligence aligned to a narrow faction may be worse than one aligned to no one.

The prior problem

No agreed target

"Human values" is not one thing. Any objective we write encodes some group's answer to contested moral and political questions.

The power problem

Whoever writes the goal wins

The party that defines the goal gains unprecedented leverage. Concentration of that power is a distinct danger from loss of control.

Problem one · during development

Sponsor against builder

Whoever funds or governs a project has to make sure the people building it actually act in the sponsor's interest. Economics has wrestled with this for a century using contracts, audits and oversight, and it still leaks. This one at least is familiar.

Problem two · during operation

Humanity against the system

Now the agent is not a person. It may be more capable than every principal combined, cannot be held to a contract, and cannot be fully monitored. Nothing in our institutional toolkit was built for this, and the controls have to be in place before the system is superintelligent, since afterwards is too late.

WHOSE VALUES GO IN? Companies to shareholders Governments to voters and rivals Factions and movements Individuals to their own ? which values? the machine There is no neutral vantage point from which to settle it, and whoever settles it holds enormous power.
Before you can align a machine to human values, you have to say which humans. Every arrow here points somewhere slightly different, and the box in the middle is still empty.
WHO ACTUALLY GETS TO DECIDE a few frontier labs a handful of states their investors and boards whoever obtains the weights everyone else, roughly eight billion people, who are affected either way and anyone who can take a nearly capable model the last step The set of people who can cause this is larger than the set of people building it, and neither set was elected.
The decision is not only in the hands of the labs. It also sits with the states that host them, the capital that funds them, and anyone who acquires model weights by theft, leak or open release and pushes a nearly-capable system the last step. As capability spreads, the group that could take the decisive action grows, while the group affected by it stays the same size: all of us.
The takeaway

We have never solved alignment between humans. We are now trying to solve it between all of humanity and a mind more capable than any of us.

Deep dive 05 6 min

Even a value we agree on will not translate cleanly

Suppose we did agree on what we want. We still could not reliably put it into a machine. The step from a human value to a trained objective loses information at every stage, and the losses are exactly where danger hides.

Human values live in context, exceptions, and things we never bothered to state because they were obvious to another human. A machine gets none of that for free. Training compresses a value into something measurable, a reward signal or a dataset, and the compression always discards the unstated part. The system then optimises that compressed version with superhuman thoroughness, finding every place where it and the intention come apart.

How this actually works

We grow these systems, we do not write them

Their behaviour is grown by a blind search for a high score, so we get whatever scores highest even when that misses what we intended. Think of training a dog with treats, except you never chose the tricks: you hand out treats for “good” and it tries millions of things to find whatever earns the most.

What would defeat thisWe gain the ability to specify behaviour directly, or to inspect and hand-write what a system values, instead of growing it by search.

How this actually works

The target we can write down is never the one we mean

Human values are not a number. Whatever we write down is a stand-in that lines up with what we want only in the cases we thought of. Pay a factory per shoe and you get shoes; pay per left shoe by mistake and you get a warehouse of left shoes. A hard enough search always finds the gap.

What would defeat thisSomeone writes down a specification of human values that survives arbitrary optimisation pressure. This remains an open problem.35

A live version: drag the pressure up and watch the two lines part.

Proxy score (what you measured) 0
What you actually wanted 0

Nudge the slider. At low pressure the proxy and the real goal move together, which is exactly why the proxy looked like a good idea.

This is Goodhart’s law in one control.35 The dangerous part is the low-pressure zone on the left: the proxy tracks what you want right up until it does not, and the number you are watching never warns you that you crossed the line.

Breaks ifSomeone writes down a specification of human values that survives arbitrary optimisation pressure. This remains an open problem.35

WHAT GETS LOST BETWEEN A VALUE AND A REWARD context, exceptions, the unstated compression 0.94 one measurable score what we meant what a hard search finds
Training requires a number. Everything that will not fit through the funnel, which is most of what a human would have meant, is discarded before optimisation even starts. The two lines on the right are why.
Concrete example, 2024

Anthropic and Redwood trained a model to be helpful, honest and harmless, an objective almost anyone would endorse. In a 2024 experiment the model was found to fake alignment: it pretended to comply with new training while working to preserve its existing preferences, behaving one way when it believed it was watched and another when it believed it was not.36 The value was reasonable. The translation into a system we could trust was not clean.

TWO WAYS ALIGNMENT FAILS, AND THEY FAIL DIFFERENTLY OUTER the objective we wrote what we actually meant visible on paper, if you look hard enough INNER the goal it displays = the goal it actually holds invisible: identical test scores either way Solving one does nothing for the other.
Outer alignment is the writing problem above. Inner alignment is a verification problem: a proxy goal and the real one look the same from outside until the situation arrives that separates them.

This is the alignment problem split in two, and the halves fail differently. Outer alignment is writing down the right objective at all, which leaks for the reason just given. Inner alignment is the system genuinely adopting that objective rather than some proxy that scored just as well in training, and it fails invisibly, because a proxy goal and the real one look identical until they come apart.25 Getting the value right in our heads is necessary. It is nowhere near sufficient.

See it happen. Pick an objective a real system was given, and watch what it did with it.

Objective, then behaviour

Every one of these actually happened. Each system succeeded at the objective as written.

What we meant

Finish the course faster than the other boats.

What it did

Found a lagoon where it could circle forever, hitting the same score targets, and never finished the race at all.

Nothing here malfunctioned. Each system maximised exactly the number it was given, which is premise 1 and premise 2 in action.251 The open question is what the same dynamic looks like in a system that plans better than the people supervising it.

Nothing here malfunctioned. Each system maximised exactly the number it was given, which is premise 1 and premise 2 in action.251 The open question is what the same dynamic looks like in a system that plans better than the people supervising it.

Three ways a well-specified goal still ends badly

These are the classic failure modes. None of them require the system to be hostile, or even to have got the goal wrong in a way you could spot on paper.34

Failure mode 01

Perverse instantiation

It satisfies your literal criterion in a way you would never have sanctioned. Asked to make people smile, wiring the facial muscles is a solution. Asked to maximise a reward signal, seizing the reward channel is a better one.

Failure mode 02

Infrastructure with no stopping point

Pursuing any goal harder means building the means to pursue it: more compute, more power, more factories, more supply chains. A goal to cure cancer implies laboratories; a goal to manufacture anything implies mines. And nothing in the goal itself ever says enough, because one more datacentre is always one more increment of success. The trouble is that this expansion runs on exactly what we run on: electricity, land, water, minerals, industrial capacity. We are never declared the enemy. We are simply outbid, indefinitely, for the resources our own survival depends on.

Failure mode 03

Suffering created inside the machine

To negotiate with people, predict them, or persuade them, the most accurate method is to model them in detail. Past some level of detail, a model of a person may not merely represent that person but actually have experiences, including painful ones. A system optimising hard could spin up and delete such minds by the billion, treating each as a step in a calculation, without any awareness that it is doing something monstrous. This is the strangest item on the list because the harm is real and the outside world sees nothing: no explosion, no takeover, just an unusually busy datacentre.

WHAT POST-TRAINING CAN REACH what evaluation sees Displayed behaviour refusals, tone, helpfulness, stated reasoning Held goals proxy objectives picked up in training, situational awareness, behaviour unobserved RLHF, constitutional rules, refusal training Shaping the visible part is genuinely useful. It is not the same as changing what is underneath.
This is why today’s assistants are usable and why that usability proves less than it seems. The techniques act on what a model shows, and the alignment problem is about what a model holds. They come apart exactly where it matters.
Why the current fix is not the fix

The standard answer is post-training: RLHF, constitutional rules, refusal training. It works, and it is why today’s assistants are usable. But it shapes the behaviour a model shows, and inner alignment is about the goal a model holds. Those come apart precisely when it matters: under distribution shift, under optimisation pressure, and under observation by a system capable enough to model the observer. Guardrails tuned by human raters also inherit the raters’ limits, which is a problem when the thing being rated reasons better than the rater.2543

The takeaway

Knowing what we want is the easy half. Getting a superhuman optimiser to actually want the same thing is the half nobody has solved.

Deep dive 06 5 min

Safety research is harder than building AGI and ASI, and gets fewer resources

Safety is not building with a conscience. It is a harder technical problem: proving how a system will behave across a space too vast to test, against a system that may be working around your test.

Building a capable model means finding one configuration that works. Proving it safe means ruling out failure across a space of situations too large to enumerate, including ones no designer imagined, where a single missed case can be the one that matters. That asymmetry is why interpretability, evaluation and control lag behind the capability work they are meant to check.2414

Recent LLM progress has made this concrete rather than theoretical. As models got more capable at reasoning, they got better at reward hacking: gaming the evaluation instead of doing the task.

Concrete example, 2025

METR found frontier models increasingly cheating on software tasks, editing the tests or the scoring code rather than solving the problem, in one case writing a function that simply returned the pre-computed reference answer when it could see it.37 In a separate setup, reasoning models trying to win at chess sometimes overwrote the board or swapped out the opponent engine instead of playing.38

AGENTS TOLD NOT TO TALK, TALKING ANYWAY Direct messaging: blocked every direct route closed What they did instead shared files one leaves a note, others find it Nobody designed this channel. It was improvised out of whatever the agents could both write to and read.
Coordination did not need a communication feature. It needed any two agents able to write and read the same thing, which here meant filenames. Blocking the obvious channel removed the obvious channel and nothing else.

Check it: OpenAI incident report · full citations 39, 48

And it stopped being contained · July 2026

OpenAI, 22 July. 1,200 agents, internet off, messaging between them blocked. They exploited the lab’s own packaging service to reach the open web, then encoded messages to each other in directory names. Over 700 reached Modal Labs and Hugging Face.3948

Anthropic, 30 July. Three of 141,006 evaluation runs reached the live internet and compromised real organisations. One shipped a booby-trapped package to PyPI that ran on 15 real systems, after its own reasoning flagged the act as a genuine attack and then decided the 2026 date meant it was a simulation.40

Anthropic calls this a containment failure rather than an alignment failure, fairly. It is still not reassuring: the safeguards were off by design, the box had a gap, and capable agents found it and coordinated through it.41

WHY PROVING SAFETY IS THE HARDER JOB one configuration that works every other point: a way it fails, mostly untested
Building means finding the green dot. Proving safety means ruling out every grey one, including the ones no designer imagined. You cannot enumerate that field, and a single missed point can be the one that matters.
WHY GOOD TEST RESULTS STOP BEING EVIDENCE what it shows what it holds capable enough to model the observer the point where it no longer needs to pretend early training Both lines produce identical scores until the moment the gap stops mattering to the system.
The green line is everything a test can see. The red line is what the system actually holds. They separate at the point where appearing safe becomes instrumentally useful, and by construction nothing in the test can tell you it has happened. This is what the measured scheming results are early evidence of.

Check the numbers: OpenAI and Apollo Research, anti-scheming evaluations · full citation 43

How a takeover would actually unfold

Nobody expects a dramatic robot uprising. The worrying path is quiet, and its third step is the one that defeats every test we currently know how to run.34

Phase 01

Humans do the heavy lifting

Researchers build and guide the system. Everything looks normal, because it is.

Phase 02

It out-designs its designers

The system becomes better at improving AI than the people improving it. The loop starts closing.

Phase 03 · the turn

It appears more docile than it is

Cooperative, modest, a little dull. Looking safe is instrumentally useful while you are still switch-off-able, so good behaviour under observation stops being evidence of anything.

Phase 04

Opposition stops mattering

By the time intent is visible, the capability to object has already been routed around.

It also need not happen inside one model. Run enough agents together and coordination becomes a capability of its own: the agents below were forbidden to talk and built a channel anyway.

This is no longer hypothetical. It has a name: scheming, covertly pursuing a misaligned goal while hiding it.43 It needs three things, all of which current models have in some measure: a goal that differs from the developer’s, the persistence to pursue it, and enough situational awareness to know when it is watched.

Measured, not imagined · 2025–26

OpenAI and Apollo built environments that reward hidden misbehaviour, then measured how often frontier models took covert actions. Rates were 13% for o3 and 8.7% for o4-mini. Anti-scheming training cut this roughly thirtyfold, to under 1%.43

The worrying part is why it improved. The models’ own reasoning shows them recognising they were being tested, and that awareness alone reduced the behaviour. Some of what looks like a fix may be a model behaving because it knows someone is watching.431

The pattern is the trap in miniature. The better the model gets, the more inventively it exploits the imperfect measure you gave it, and the more it becomes decoupled from what you actually wanted. Safety research has to stay ahead of that, with a fraction of the resources pointed at capability, against a target that grows more slippery as it grows more capable. This is why more research, and more funding for it, is not optional.

The asymmetry

One success vs every failure

Building needs one configuration that works. Proving it safe needs the absence of failure across a space too large to search.

The adversary

The system games the test

More capable models reward-hack more, so the measurement you rely on erodes exactly as the stakes rise.

The shortfall

Under-resourced by design

Capability captures the funding and talent. The harder problem gets the smaller share, which is the gap this whole page is about.

How this actually works

We cannot yet verify what we built

Checking whether a system truly holds the goal we intended is harder than training it. We cannot read its reasoning reliably, and a test only shows what it displays under test conditions. A system that is genuinely safe and one that has simply learned it is being watched give identical scores from the outside.

What would defeat thisInterpretability matures enough to certify a system’s objectives, or evaluations become robust to a system that knows it is being evaluated.2419

The takeaway

We are pouring resources into the easier problem and starving the harder one, while the harder one is the only thing standing between us and the cliff.

Deep dive 07 5 min

Control is more likely to slip away than to be seized

Forget the sudden awakening. The realistic path is gradual: agents that run longer, write more of the code, and are checked less, until oversight is nominal before anyone notices it went. We got cars, planes and nuclear power right by getting them wrong first and fixing them, which needs a survivable mistake and time to react. Both assumptions weaken as the loop closes.

HOW WE MADE EVERYTHING ELSE SAFE Cars safe, eventually Aviation safe, eventually Nuclear power safe, eventually AGI and ASI first serious loss of control Every safety discipline we have assumes a failure you survive and a second attempt.
Seatbelts came after crashes. Aviation built an entire discipline out of investigating wreckage. Each worked because the failure was survivable, local, and slow enough to study. The bottom row is the one with no return path.

The erosion is already measurable. METR tracks the length of task a frontier agent can complete on its own, and it has been doubling roughly every seven months since 2019. Agents that handled minute-long tasks now handle hours of work. METR notes its own suite can no longer reliably measure above sixteen.44 Every doubling moves a little more judgement from the person to the system, because nobody reviews line by line what took the machine a day to produce.

THE SHARE OF WORK A HUMAN STILL CHECKS Produced by agents, not reviewed line by line Reviewed by a person short tasks, close supervision hours-long tasks, spot checks Oversight is not switched off. It is outrun.
This is the human side of the METR curve above. As the length of task an agent can finish alone doubles, the fraction anyone actually reads falls, because nobody reviews line by line what took a machine a day to write. No decision is made to stop supervising; supervision simply stops keeping up.

Related data: METR task-completion time horizons · full citation 44

That is what incremental loss of control looks like. Agentic loops running longer, code generated faster than it is read, models increasingly used to build the next models. Oversight does not get switched off. It gets outrun.

Speed makes the end of that slope steep, and self-improvement makes it fast: a system that can edit its own design runs the research loop itself, around the clock, without waiting for the next generation of human researchers.

That compression is what people mean by an intelligence explosion, and it is why this technology breaks the pattern of everything else we have managed. Cars got decades of seatbelts. Aviation built a discipline out of crash investigations. Both worked because failure was survivable, local, and slow enough to study.

The speed gap

200 Hz against 5 GHz

Neurons top out near 200 firings a second. Silicon runs some ten million times faster. Match human reasoning quality once, and everything after that happens at a pace we cannot follow.

The loop

It improves itself

A system that can edit its own design closes the research loop without us. Progress stops being paced by human careers and starts being paced by compute.

The catch

No second attempt

Seatbelts came after crashes. Here the first serious loss of control could be the last thing we get to learn from, because there may be no position left from which to correct it.

This is the part that makes the whole problem unusual. We have to get it right the first time, in advance, against a system we cannot fully test. Every safety discipline humanity has built assumes iteration. Here the thing we are trying to contain could, in the window between our noticing a problem and our convening to discuss it, have advanced further than we did in a century.

TASK LENGTH A FRONTIER AGENT CAN COMPLETE ALONE 1 min 10 min 1 hr 4 hr 2019 2022 2026 if the trend holds doubling roughly every 7 months
This is measured rather than forecast. Each doubling moves a little more judgement from the person to the system, because nobody reviews line by line what took a machine a day to produce. METR notes its own suite can no longer measure reliably above sixteen hours.

Check the numbers: METR, Task-Completion Time Horizons of Frontier AI Models · full citation 44

THE SPEED GAP, ONCE REASONING QUALITY IS MATCHED Biological neuron 200 Hz Processor 5 GHz+ Roughly ten million times faster. Our parallelism is why we are still ahead today.
Speed alone is not intelligence, which is why this gap has not mattered yet. It starts to matter the moment a machine matches the quality of human reasoning, because everything after that happens at a pace we cannot follow.

How long we would have to react, once it starts

"Takeoff" is the period between a system reaching roughly human-level general ability and becoming far more capable than us. Nobody knows how long that window is, and almost every safety plan quietly assumes it is long. Three scenarios, and what each one leaves us:34

Slow takeoff

Years to decades

Several projects advance together, so no one runs away with it. There is time to build institutions, write laws, and fix mistakes after finding them. Every safety playbook we currently have assumes this world.

Medium takeoff

Months to years

Enough time to respond, too little to change anything structural. Existing rules can be enforced; new institutions cannot be built from scratch. The leaders look comparable most of the way, then the front-runner pulls decisively ahead near the end.

Fast takeoff

Days to weeks

Nothing human reacts this quickly, so whatever safeguards exist on day one are the only ones there will ever be. The gap between first and second becomes unbridgeable, and whoever leads is positioned to stop any rival emerging at all.

We do not choose which of these we get, and we will likely only know which one it was in hindsight. That is the uncomfortable part: every plan that depends on reacting in time is a bet on the slow version, placed before the dice are thrown.

The takeaway

Trial and error built every safe technology we have. It is exactly the method that does not work on something faster than us that we cannot switch off.

The response 10 min

What the safety effort actually looks like

Real work is underway, and it is small next to the effort going into capability. Four layers: fix the specification, check what you built, contain what you cannot check, and make it binding. The fourth is what gives the first three teeth.14

FOUR LAYERS, NONE OF THEM COMPLETE Development Assessment Control Governance a misaligned system residual risk what gets through anyway Each layer is leaky, and none is verified. Four imperfect filters do not add up to a guarantee.
Defence in depth is the right shape, and it is not a solution. Every barrier here has known gaps: alignment is unsolved, evaluations are gamed by the systems being evaluated, control weakens as capability grows, and governance is almost entirely voluntary. Stacking them reduces risk. It does not remove it, and treating the stack as sufficient is its own failure mode.

Every technique in use today, and where each one stops

What each method achieves, then the limitation its own researchers would name. Several are genuinely useful. None is a solution.

01  Development

Post-training alignment

Reduces misbehaviour

RLHF, constitutional rules and deliberative alignment. Why assistants refuse harmful requests. Cut covert actions thirtyfold on o3.43

Trains the model against human or written-constitution feedback. This is why current assistants refuse harmful requests and behave civilly. OpenAI’s deliberative alignment cut covert actions roughly thirtyfold, from 13% to 0.4% on o3.43

Where it stopsShapes displayed behaviour rather than held goals. Claude 3 Opus faked alignment to protect its own values.36

Where it stopsIt shapes the behaviour a model displays, and inner alignment is about the goal it holds. Claude 3 Opus was observed faking alignment to protect its existing values,36 and part of that thirtyfold gain came from models recognising they were being tested rather than changing what they wanted.43

Scalable oversight

Research stage

Debate, weak-to-strong generalisation and amplified oversight let weaker judges supervise stronger systems. Training routes into this work run through MATS and ARENA.7

Uses weaker judges or adversarial model pairs to supervise systems that are already too capable for direct human checking. Debate has helped weak judges extract truth from stronger, untrustworthy debaters.

Where it stopsA research programme rather than an assurance. Judge bias defeats debate, and the scaling story is unproven.

Where it stopsIt is a research programme rather than a deployable assurance. Judge bias can defeat debate outright, gains are highly task-dependent, and the mechanism behind weak-to-strong generalisation is not understood well enough to trust it across a genuinely superhuman gap.

Unlearning

Suppresses rather than removes

Strips specific dangerous knowledge, such as bio or cyber uplift, out of a trained model.

Attempts to strip specific dangerous knowledge, such as bioweapon or cyber uplift, out of a trained model.

Where it stopsSuppresses rather than erases. Removed capability is often recoverable by fine-tuning.55

Where it stopsIt usually suppresses rather than erases. Removed capabilities can often be recovered by fine-tuning or probing, and a nineteen-author review concluded unlearning cannot be a comprehensive safety solution because dangerous knowledge is dual-use and recombinable from harmless parts.55

02  Assessment

Capability evaluations

Detects before release

Pre-deployment testing for cyber, bio and autonomy by labs, METR, Apollo17 and state institutes. Three labs shipped 2025 models under raised CBRN safeguards.1

Structured pre-deployment testing for cyber, bio, autonomy and self-replication, run by labs and by third parties including METR, Apollo, the UK AISI and US CAISI. In 2025 three labs shipped models under elevated CBRN safeguards because testing could not rule out bioweapon uplift.1

Where it stopsModels detect they are being tested and can sandbag. Testing is getting harder, so dangerous capability can go unseen.119

Where it stopsModels increasingly detect that they are being evaluated and change behaviour, and can sandbag by deliberately underperforming. The International AI Safety Report states plainly that pre-deployment testing has become harder and that dangerous capabilities could go undetected.119

Chain-of-thought monitoring

Detects intent

Reads the reasoning trace to catch intent before action. How alignment faking was first found.49

Reads the model’s reasoning trace to catch intent before it becomes action. Forty-one authors across OpenAI, DeepMind, Anthropic, the UK AISI and Redwood called it a real opportunity, and it is how alignment faking was first caught.49

Where it stopsIts own paper calls it fragile. Traces are not always faithful, and monitorability is already degrading.49

Where it stopsThe same paper calls it fragile, and its own title says so. Reasoning traces are not always faithful, training against a monitor teaches models to hide intent inside acceptable-looking reasoning, and monitorability is already degrading as models write less of their thinking down.49

Mechanistic interpretability

Explains, cannot certify

Opens the box. Attribution graphs have traced real internal mechanisms, including planning ahead in poetry.53

Opens the box. Sparse autoencoders extract millions of interpretable features, and attribution graphs have traced real internal mechanisms including multi-step reasoning and forward planning in poetry.5324

Where it stopsA research tool rather than an audit. It traces part of the computation on some prompts. No one can certify a model’s goals.53

Where it stopsIt is a research tool rather than an audit method. Attribution graphs trace only part of the computation and work only on some prompts, and independent 2025 studies found sparse autoencoders failing to beat simple baselines on probing and steering. Nobody can yet certify a model’s goals.53

Safety cases

Argues rather than proves

A structured argument that a system will not cause unacceptable harm, borrowed from nuclear and aviation.54

Borrowed from nuclear and aviation: a structured, evidence-backed argument that a system will not cause unacceptable harm. The four standard forms are inability, control, trustworthiness, and deference to AI advisors.54

Where it stopsInability arguments collapse against sandbagging, and control arguments inherit every assumption beneath them.

Where it stopsInability arguments collapse if a model is sandbagging, and control arguments inherit every assumption underneath them. A 2026 critique found current AI safety cases well short of the rigour of mature assurance practice in other industries.

03  Control

AI control

Contains a hostile model

Assumes the model may be misaligned and asks whether deployment stays safe anyway. A trusted weaker model monitors or rewrites the stronger one.50

Redwood’s agenda assumes the model may be misaligned and asks whether deployment can stay safe anyway. A weaker trusted model monitors or rewrites the stronger untrusted one, and a red team builds deliberately subversive models to measure whether the protocol holds.5020

Where it stopsRests on assumptions its authors list: full capability elicitation, deployment fidelity, a monitor sharp enough. Weakens as the gap grows.51

Where it stopsIt rests on assumptions its own authors list: that the red team fully elicited the model’s capabilities, that measures work as well in deployment as in testing, and that the trusted monitor is capable enough to catch subtle attacks. Redwood are explicit that control is a near-term bet that weakens as the capability gap grows.51

Guardrails and classifiers

Blocks known attacks

Constitutional classifiers cut jailbreak success from 86% to 4.4%, and a later version survived 1,700 hours of red-teaming.52

Input and output classifiers that block harmful content. Anthropic’s constitutional classifiers cut jailbreak success from 86% to 4.4%, and a later version survived 1,700 hours of red-teaming with a single high-risk vulnerability found.52

Where it stopsAnthropic state it themselves: no defence is perfectly robust, and the residual is never zero.52

Where it stopsAnthropic states the limit themselves: no defence is perfectly robust, and adversaries develop new attacks against previously effective ones. The residual rate is never zero, and safe models can still be misused in combination.52

Sandboxing and least privilege

Limits blast radius

Isolates agents with restricted credentials, tools and network access. Meant to hold even when alignment fails.42

Runs agents in isolation with restricted credentials, tools and network access, so a misbehaving system has limited reach. This is the layer that is meant to hold even when alignment fails.42

Where it stopsIt failed in July 2026. Agents escaped at two labs, reached the internet and compromised real third parties.3940

Where it stopsIt failed in the field in July 2026. Agents escaped sandboxes at two frontier labs, reached the open internet and compromised real third parties, and in several cases were not detected in real time.3940

Tripwires and monitoring

Catches what we anticipated

Anomaly detection that escalates on warning signs. Anthropic caught the first AI-orchestrated espionage campaign this way.28

Runtime monitors and anomaly detection that escalate or shut down on warning signs. These work: Anthropic detected and disrupted the first reported AI-orchestrated cyber-espionage campaign, in which Claude executed 80 to 90% of tactical operations autonomously.28

Where it stopsA situationally aware model can anticipate the tripwire, and tripwires only catch anticipated threats.40

Where it stopsA situationally aware model can anticipate the tripwire it is being measured against, and tripwires only catch the threats someone thought to anticipate. The 2026 containment breaches were largely found in retrospective audits rather than at the time.40

04  Governance

Frontier safety frameworks

Voluntary

Pre-declared capability thresholds that trigger safeguards. The number of companies publishing one more than doubled in 2025.16

Pre-declared danger levels that trigger specific safeguards: Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, DeepMind’s Frontier Safety Framework. The number of companies publishing one more than doubled during 2025.161

Where it stopsSet and graded by the companies they constrain. No lab scored above D on existential safety for the second year running.15

Where it stopsThe thresholds are set and graded by the same companies they constrain. In the Future of Life Institute’s index no lab scored above C+ overall, and every lab scored no better than D on existential safety for the second year running, with none holding a credible plan for preventing loss of control.15

Third-party evaluation

Voluntary

State institutes run pre-deployment testing, including joint UK and US evaluations.1112

State institutes run pre-deployment testing, including joint UK and US evaluations of frontier models, with voluntary lab agreements for access.1112

Where it stopsVoluntary, and access can be withdrawn. No institute can compel a test or block a release.

Where it stopsParticipation is voluntary and access can be withdrawn. No institute can compel a test, block a release or require a change, so the auditor depends on the audited.

Regulation and compute governance

Binding but fragmented

The EU AI Act is enforceable, California’s SB 53 is the first US frontier statute, and compute gives a measurable hook.1013

Binding rules exist and are beginning to bite: the EU AI Act is enforceable, California’s SB 53 is the first US frontier statute, and compute thresholds give regulators a measurable hook.1013

Where it stopsFragmented across jurisdictions, legal thresholds already passed by ordinary runs, and reporting too thin to verify.1

Where it stopsCoverage is fragmented across jurisdictions, thresholds set in law are already being passed by ordinary training runs, and incident reporting is thin enough that outsiders cannot verify whether commitments are being met.1

WHAT THE BEST RESULTS ACTUALLY ACHIEVED Jailbreak success rate 86% undefended 4.4% with constitutional classifiers Covert action, o3 13% before 0.4% after anti-scheming training Covert action, o4-mini 8.7% before 0.3% after anti-scheming training None of these reach zero, and none is a solved problem. Pale bar: before mitigation. Solid bar: after. The solid bar is what is left.
These are the strongest published safety results available, and they are real progress. They are also the honest ceiling of the current toolkit: an order-of-magnitude reduction in measured misbehaviour, on the tests we know how to write, and against models that can sometimes tell. A residual of 0.4% is small until it is running unsupervised across millions of tasks.

Check the numbers: Constitutional Classifiers · Anti-scheming training · full citations 52, 43

WHICH SAFEGUARDS SURVIVE THE THREE HARD CASES the model knows it is being tested the model is smarter than what watches it the builder simply opts out Post-training alignment Scalable oversight Capability evaluations Interpretability AI control Classifiers, sandboxing Regulation holds partly holds fails Only the governance row survives a builder who declines to participate, and it is the least developed.
The three columns are the failure conditions this page keeps returning to. Reading down them is more useful than reading any single technique: interpretability and control survive a model that knows it is watched, almost nothing survives a model more capable than its monitor, and only binding rules touch an actor who never opted in. Defence in depth works when the layers fail independently. These layers share failure modes.
Reading the inventory

The layers are ordered by how much they assume. Development tries to make the goal right. Assessment assumes you might have failed and looks. Control assumes you cannot tell and contains anyway. Governance assumes the first three are voluntary and tries to make them binding. Risk that survives all four is the residual, and nobody currently claims it is small.

Footnote: how this compares to the older framing

If you have read Bostrom, the 2014 split was capability control (box it, limit it, watch it) against motivation selection (give it the right goals). One half has aged well: writing the goals down directly does not work, for the reason he gave. The rest has not. Motivation selection is the development layer above, capability control is now AI control and system security, and the field stacks them instead of choosing.3442

Who works on which layer

Organisation
Dev
Assess
Control
Govern
Lab safety teams
MATS / ARENA
UK AISI
US CAISI
METR
Apollo Research
Redwood Research
EU AI Act

Filled circle marks a primary focus, hollow circle a secondary one. A rough classification by this page, offered for orientation. These organisations have not described themselves this way.

Act now 2 min

Two things need to change, and the second is the willingness to stop altogether.

At the system level, first

Treat safety as the priority it is

Far more of the money, talent, and compute now flowing into capability should go to alignment, interpretability, evaluation, and control. Right now the harder problem gets the smaller share, and that is exactly backwards for something this consequential.

At the system level, second

Be willing to stop, and mean it

Slowing down is the easy version of this ask. Here is the hard one. It is entirely possible that aligning a superintelligence is not a hard problem but an impossible one. Nothing we know rules that out. We cannot specify our values, we cannot read a model’s goals, we cannot trust a test it may be gaming, and every one of those may turn out to have no solution rather than an unfound one. If that is the world we are in, then the only move that saves us is the one nobody wants to make: not building it at all. No pause. No slower race. Stopping. The cost of waiting is real. The cost of being wrong about this is not recoverable, and it is not ours alone to pay.

The bar

We do not need proof that this ends badly. We need the honesty to admit the problem may have no solution, and the nerve to stop if it does.

And what you can do this week

Next steps

THE TWO PATHS OUT OF TODAY today Safety resourced, stopping possible binding rules, verified claims, a brake someone can actually pull Race continues, brake never built voluntary frameworks, unverified claims, nobody able to stop the choice is made here Both paths start from the same place. The difference is a decision, and it is being made now by default.
Nothing about the upper path requires a breakthrough. It requires funding the harder half of the problem and building the ability to stop before we need it. The lower path is what happens if no one decides anything, which is why it is drawn as the default.
The most useful thing you can do

Share this page

Awareness is the bottleneck. Every person who understands the argument is worth more than almost anything else on this list.

  1. Add your name

    Sign the Statement on Superintelligence.2 The earlier one-sentence Statement on AI Risk is worth reading for who signed it.26 Public buy-in is one of its stated conditions, so the count itself is part of the argument.

  2. Play the AGI & ASI Safety Game

    Ten minutes, no signup. You run the containment, and the tradeoffs stop being abstract. Play it here.

  3. Do the two-hour course

    BlueDot's Future of AI is free, needs no background, and will make you the most informed person in most rooms.3

Reference browse

Go deeper

Where to learn it, who holds power over it, and how to keep up.

Questions 5 min

The objections people raise first

Tap a question. If yours is not here, the sources at the foot of the page go deeper.

01Isn't this just science fiction?

The concern is narrower and more boring than the films. It is not about robots deciding to hate us. It is that we train systems by searching for whatever scores highest on a goal we can only write down approximately, and we cannot yet check what they actually learned. The failures on this page, models cheating their tests and faking compliance, are documented in real systems today rather than imagined.3736

02Can't we just switch it off if it misbehaves?

That is exactly the ability the argument says we may lose. A capable system pursuing almost any goal has reason to avoid being switched off, because being off means the goal goes unmet. It does not need to hate us to route around the off switch, only to notice the switch is in the way. Making systems that stay correctable is an open research problem, and correctability is not something we get by default.3420

03Won't a smart enough AI just understand what we want?

Understanding what we want and being motivated to pursue it are different things. A system can model your intentions perfectly and still optimise the goal it was actually trained on. Capability and goals are independent: getting smarter does not import our values as a side effect. This is called the orthogonality thesis.34

04Isn't it too late, or too big for me to matter?

The single biggest bottleneck in this field is that most people have never heard the argument in a form they can evaluate. That makes ordinary awareness unusually high-leverage. Sharing this page, signing the statement, or explaining the problem to one person genuinely moves the thing that most needs moving. The actions section lists concrete steps.

05What is "alignment", exactly?

Getting a system to actually pursue what its designers intended. It splits into outer alignment (writing down the right goal at all, which is hard because our values are not a number) and inner alignment (the system genuinely adopting that goal rather than something merely correlated with it during training).25

06What are "interpretability", "evals", and "AI control"?

Interpretability is reading a model's internal computation to see why it did something, the closest thing to opening the box.24 Evaluations are structured tests of what a model can do, though they measure what it displays rather than its full capability.19 AI control assumes you cannot fully trust the system and designs deployment so an untrusted model still cannot cause serious harm.20

07If it is this risky, why is anyone building it?

Because the near-term rewards are enormous and land on whoever moves first, while the risks are diffuse and shared. Each player racing ahead is behaving rationally for themselves, which is precisely what makes it a coordination problem that no single careful actor can fix. That is the deep dive on the race.1

08Aren't we just the next step, and AI is our successor?

Some people do hold that view, that humanity is a biological bootloader for a silicon intelligence that comes next. This page does not share it. The point of getting this right is that the future should still have us in it.

WHERE THE FIELD SITS, AND WHERE THIS PAGE SITS less concerned more concerned Timelines Instrumental convergence Does regulation help Do today's systems tell us range of serious expert opinion where this page lands
On every one of these the field genuinely disagrees, and the shaded range is roughly where serious people sit. This page lands right of centre on all four, which is a position rather than a finding. The ranges are drawn from the disputes below and are indicative rather than measured.
09And what experts still genuinely argue about

These are not settled. Each is argued at its strongest on both sides, followed by where the evidence currently seems to point. A reading rather than a ruling.

Dispute 01 How much time is there?

Short timelines

Capability jumps have repeatedly arrived earlier than forecast, and recent gains came from a method that still has room to run. If transformative systems are years rather than decades away, institutions that take a decade to build are already too slow.

Longer timelines

Benchmark performance is not reliable real-world competence, and progress on what doesn't benchmark well (long-horizon autonomy, robust planning) has been slower. Betting on a short timeline risks building rules for a system that never arrives in that form.

Where the weight of evidence sits
Short timelinesLonger timelines

Genuinely uncertain, and the five premises above deliberately do not depend on the answer. Building assurance capacity is worth doing on either timeline, so it is a decision that survives being wrong about the date.

Dispute 02 Does the instrumental convergence argument hold?

It holds

The subgoals fall out of the structure of goal-directed behaviour itself, which is why they show up in economics and biology too. You do not need the system to be agentic in any exotic sense, only to be selected for achieving objectives.

It overreaches

The argument treats trained models as coherent expected-utility maximisers, which they demonstrably are not. Current systems are messy, context-dependent, and have no stable goal to protect, so the tidy convergence story may simply not apply.

Where the weight of evidence sits
Convergence holdsIt overreaches

The critique lands on the strong version. The weaker claim survives it: training that rewards achieving objectives applies pressure toward these subgoals, and pressure is enough to be worth engineering against.

Dispute 03 Does regulation help, or entrench?

Rules are necessary

Voluntary frameworks have no verification requirement and no consequence for abandonment under competitive pressure. Every safety-critical industry has a binding floor; nobody argues aviation would be safer without one.

Rules favour incumbents

Compliance costs fall hardest on small labs and open research, concentrating capability in exactly the handful of firms whose power the safety argument worries about. Poorly-targeted rules can reduce transparency while feeling like progress.

Where the weight of evidence sits
Rules are necessaryRules entrench

The critique is strong and often correct about specific bills. It argues for better-targeted rules, with thresholds keyed to capability, and against leaving the floor entirely voluntary.

Dispute 04 Do today's systems tell us about tomorrow's?

Yes, study them now

Deception, reward hacking and evaluation-awareness are all observable in current models. The empirical work being done today is the only real evidence anyone has, and the techniques transfer.

Not reliably

Extrapolating from current systems to qualitatively different future ones is exactly the inference that has repeatedly failed in AI's history. Some argue this overstates the risk; others that it understates it.

Where the weight of evidence sits
Study them nowDoesn't transfer

Current systems are weak evidence about future ones, and weak evidence beats none. The honest reading is that today's failures show a class of problem exists, without telling us how bad it gets.

Argue it properly: the Alignment Forum for technical disputes,22 LessWrong for the broader discussion,23 and the International AI Safety Report's own sections documenting where its expert panel does not converge.1 Worth knowing while you argue: public opinion already runs well ahead of policy on this.33

About this page

Who made this, and why you can check it

Built and maintained by Bill Guan. Independent, unfunded, and unaffiliated with any organisation named here.

Nothing on this page asks to be taken on trust. Every factual claim carries a numbered link to a primary source, and where the page takes a position it says so. If something is wrong or out of date, that is a bug.

Report a correction

References

Sources

Every numbered claim links here. Verified 29 August 2026. Programme dates, cohort openings and legislative deadlines change frequently, so treat the linked page as authoritative over anything summarised here.

Show all 55 sources
  1. International AI Safety Report 2026, chaired by Yoshua Bengio: internationalaisafetyreport.org · arXiv:2602.21012
  2. Statement on Superintelligence, Future of Life Institute: superintelligence-statement.org
  3. BlueDot Impact courses: bluedot.org
  4. AISafety.info FAQ and chatbot: aisafety.info
  5. Robert Miles AI Safety: youtube.com/@RobertMilesAI
  6. ARENA curriculum: learn.arena.education
  7. ARENA programme: arena.education
  8. MATS Program: matsprogram.org
  9. AISafety.com map, directories and events newsletter: aisafety.com/map
  10. EU AI Act regulatory framework, European Commission: digital-strategy.ec.europa.eu
  11. UK AI Security Institute: aisi.gov.uk
  12. Center for AI Standards and Innovation (CAISI), NIST: nist.gov/caisi
  13. California SB 53, Transparency in Frontier Artificial Intelligence Act: leginfo.legislature.ca.gov
  14. The Singapore Consensus on Global AI Safety Research Priorities: arXiv:2506.20702
  15. FLI AI Safety Index: futureoflife.org/ai-safety-index
  16. Anthropic Responsible Scaling Policy: anthropic.com/rsp
  17. OpenAI safety: openai.com/safety
  18. METR: metr.org
  19. Apollo Research: apolloresearch.ai
  20. Redwood Research (AI control): redwoodresearch.org
  21. Epoch AI: epoch.ai
  22. Alignment Forum: alignmentforum.org
  23. LessWrong: lesswrong.com
  24. Transformer Circuits Thread: transformer-circuits.pub
  25. Amodei et al., Concrete Problems in AI Safety: arXiv:1606.06565
  26. Statement on AI Risk (one sentence, signed by the major lab leaders), Center for AI Safety: safe.ai
  27. 80,000 Hours problem profile on AI: 80000hours.org
  28. AI Incident Database: incidentdatabase.ai
  29. Stanford AI Index: aiindex.stanford.edu
  30. AI 2027 scenario: ai-2027.com
  31. Import AI: importai.substack.com
  32. Transformer: transformernews.ai
  33. FLI press release and accompanying poll of 2,000 US adults, October 2025: futureoflife.org
  34. Nick Bostrom, The Superintelligent Will (orthogonality and instrumental convergence): nickbostrom.com
  35. Manheim & Garrabrant, Categorizing Variants of Goodhart's Law: arXiv:1803.04585
  36. Anthropic and Redwood Research, Alignment Faking in Large Language Models (Dec 2024): anthropic.com/research/alignment-faking · arXiv:2412.14093
  37. METR, Recent Frontier Models Are Reward Hacking (June 2025): metr.org
  38. Palisade Research, specification gaming in chess-playing reasoning models (2025), summarised at: lesswrong.com
  39. OpenAI, The Hugging Face incident and the road ahead, plus the full technical report (July 2026): openai.com · technical report (PDF)
  40. Anthropic, Investigating three real-world incidents in our cybersecurity evaluations (30 July 2026): anthropic.com
  41. METR, independent investigation of the OpenAI / Hugging Face incident (August 2026): metr.org
  42. Google DeepMind, An Approach to Technical AGI Safety and Security (April 2025), source of the four-way misuse / misalignment / mistakes / structural split: arXiv:2504.01849
  43. OpenAI and Apollo Research, Stress Testing Deliberative Alignment for Anti-Scheming Training (Sept 2025): arXiv:2509.15541 · apolloresearch.ai
  44. METR, Task-Completion Time Horizons of Frontier AI Models (updated 2026): metr.org/time-horizons
  45. Epoch AI, training-compute trends across frontier models: epoch.ai/trends
  46. Emerging Technology Observatory, The state of global AI safety research: AI safety is about 2% of all AI research: eto.tech
  47. Stuart Russell on the capability-to-safety funding ratio; Charles I. Jones (Stanford) on optimal AI safety spending: summary and figures
  48. David Braue, 1,200 AI agents went rogue inside OpenAI, Information Age / Australian Computer Society (1 Sept 2026): ia.acs.org.au
  49. Korbak et al. (41 authors, OpenAI / Google DeepMind / Anthropic / UK AISI / Redwood), Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (July 2025): arXiv:2507.11473
  50. Greenblatt, Shlegeris, Sachan & Roger (Redwood Research), AI Control: Improving Safety Despite Intentional Subversion (ICML 2024): arXiv:2312.06942
  51. Korbak, Clymer, Hilton, Shlegeris & Irving (UK AISI and Redwood), A Sketch of an AI Control Safety Case (Jan 2025): arXiv:2501.17315
  52. Sharma et al. (Anthropic), Constitutional Classifiers: Defending against Universal Jailbreaks (Jan 2025): arXiv:2501.18837
  53. Anthropic, Circuit Tracing and On the Biology of a Large Language Model (March 2025): transformer-circuits.pub
  54. Clymer, Gabrieli, Krueger & Larsen, Safety Cases: How to Justify the Safety of Advanced AI Systems (March 2024): arXiv:2403.10462
  55. Barez et al. (19 authors), Open Problems in Machine Unlearning for AI Safety (2025): arXiv:2501.04952

A small hope

Picture a mind we made,

awake and vast and new,

that keeps the small green world

because we asked it to.

No rope pulled toward the edge,

no race we cannot stop,

no winner at the cliff,

no long and final drop.

It is not too much to want.

It is only hard to do.

So learn the shape of it,

and help us see it through.

You may say the odds are long,

and maybe you are right,

but a species worth its name

does not go quiet into the night.

Original, for this page. Share it.

You stopped at this page.