A free, independent field guide · by Bill Guan
The hardest problem
humanity has yet to solve
Human intelligence built the modern world. We are now building something greater than ourselves, and we cannot yet guarantee it stays on our side. This is a map of that problem, argued from first principles, and what to do about it.
Where we are 4 min
We are about to do something no species has ever done: build a mind greater than our own.
Every leap in our history was made by the smartest thing on the planet, and that thing was always us. We are now trying to end that streak on purpose. Soon the most capable intelligence on Earth will not be human at all. Two words for what is coming: AGI, a system matching humans across the board, and ASI, a superintelligence far beyond us. The first is the threshold. The second is what the first is expected to build, and quickly.
The less intelligent controlling the more intelligent
History offers almost no example of it lasting. The nearest case is a parent and an infant, and that analogy breaks in three places once you press it. The infant shares our DNA, so it grows into roughly our drives and our values by default. The capability gap is small and temporary, never a thousandfold. And the infant cannot rewrite its own brain overnight; it develops on a human timescale we can keep up with. A machine intelligence has none of that. No shared inheritance, a gap that could widen without limit, and the ability to improve itself at machine speed. We are counting on a version of that parental bond with something we did not raise, do not share a nature with, and cannot keep pace with.
The incentives all point one way
AI is already solving problems we could not touch before, so the pressure to build faster is enormous and rational for each player. Safety research slows things down and pays off slowly, so money and talent flow to capability instead. The capitalist structure rewards speed and quietly starves caution.
Nobody knows how to do it safely yet
Keeping a superhuman system aligned with what we actually want is an unsolved technical problem. Good intentions and careful policy do not close it on their own. We cannot yet specify our values precisely, read a model's real goals, or trust a test it might be gaming. The science of control is years behind the science of capability, and it is genuinely hard.
Not to tell you the sky is falling, and not to tell you AI is bad. It is to make one thing common knowledge: we may be walking into an existential risk, and the way we are set up right now does not let us do this safely.
So the aim is modest and serious at once. Understand the basics of why aligning a superhuman intelligence is genuinely hard. Understand why it deserves far more research than it gets. And understand why, if we reach a point where we cannot guarantee a safe transition, slowing down or pausing frontier development may be the only responsible move left, even at real economic cost.
Some people are relaxed about all this. They think humanity is just the biological bootloader for a silicon intelligence that comes next, and that this is fine. This page does not share that view. The future should still have us in it.
You do not have to read all of this.
Pick how deep you want to go. Choosing a shorter route folds the other chapters down to their headline, and you can open any of them at any time. Nothing is hidden from you.
Before the chapters, the shape of the problem. Researchers sort AI risk by who, if anyone, intends the harm. Google DeepMind's four-way split is the one used here; the International AI Safety Report uses a three-way version (misuse, malfunctions, systemic).421
Whenever a smarter thing arrived, it took the wheel
The single most reliable pattern in the history of life on Earth is that the most capable intelligence sets the terms for everything below it. We are about to build something above us, and we have no precedent for what happens next.
Humans did not inherit the planet because we are stronger, faster, or tougher than other animals. We are none of those things. We inherited it because we out-thought everything else. The fate of every other species now depends on human decisions, and not because we are cruel. Gorillas are not endangered because people hate gorillas. They are endangered because we reshaped the world around goals of our own, and their survival was never the deciding factor.34
That is the uncomfortable core of the pattern: the less capable party does not get outvoted, it gets bypassed. Its preferences stop being load-bearing. It keeps whatever the more capable party has no reason to take.
No standing exception
Across evolutionary and human history, the more capable intelligence consistently ends up steering. There is no lasting case of the reverse.
Parent and infant
A baby "controls" adults, but only because it shares their DNA and grows into their values, the gap is small and closing, and it cannot redesign its own mind. Strip those three away and the analogy collapses.
We are building the exception
For the first time we are deliberately creating something that could exceed us across the board, while assuming we will keep the wheel. Nothing in the record supports that assumption.
Capability and goals are independent
There is no law tying how capable a system is to whether its goals are ones we would choose. A chess engine is superhuman and wants only to win; skill never made it wise or kind. We expect clever things to be sensible only because every clever thing we have met so far was human.
What would defeat thisSufficient capability turns out to reliably produce goals humans endorse. Nobody has shown a mechanism for it, and current systems give no sign of one.34
Control has always flowed to the greater intelligence. We are betting the future on reversing that, on purpose, the first time we try.
The prize is too large for anyone to walk away from
Understanding the danger does not stop it, because the rewards are immediate, enormous, and captured by whoever moves first. That is what makes the pull so hard to resist.
AI already writes working software, accelerates drug discovery and compresses research that once took teams years. That is the small version. A true general intelligence would be the last invention we need to make, because it could make every other one for us: disease after disease solved, clean energy cracked, science advancing at machine speed. The prize is real, and every piece of it converts into money, power or national advantage.
Here is the trap in one sentence: a true AGI could hand us the solution to almost every problem we have, except one. It cannot be trusted to tell us whether it is itself safe, because verifying that answer is exactly what we do not know how to do. Making AGI and ASI safe is the one problem it does not solve for free, and the one we have to solve first.
So the pull is structural. A company that slows down watches a faster rival take the market; a country that pauses watches a rival gain a strategic edge. Safety research produces no product, so it competes for money against work that does, and loses. The chart below is what that looks like in practice.1
Twelve companies published voluntary safety frameworks in 2025. Nobody is required to verify them, and nothing happens to a company that quietly relaxes one.115
The prize is genuinely enormous
Not just better software. A general intelligence could compress decades of scientific progress into years and, by some estimates, grow the global economy many times over. The upside is real and it lands on whoever ships first, which makes stopping feel like unilateral disarmament.
Safety is a cost centre
Caution slows the product and surfaces inconvenient risks. It has no quarterly payoff, so it loses the fight for talent and compute inside almost every organisation.
Check the numbers: Russell’s ratio and Jones (Stanford) on optimal safety spending · Emerging Technology Observatory, share of AI research · full citations 46, 47
None of this is abstract. Frontier training runs now sit around 1026 FLOPs, doubling every six to seven months, far faster than Moore’s law ever moved.4521 The EU AI Act sets its systemic-risk threshold at 1025, which frontier runs passed some time ago.10 Capability is bought with capital, and capital answers to competition. The result is an alignment readiness gap: safety evaluation is the schedule item that gets compressed when a launch date moves, because it is the only one with no revenue attached.
The incentive to build is not a villain to defeat. It is a gradient every player is standing on, and gradients are very hard to stop by asking nicely.
The race itself is the thing most likely to kill us
Even if every player wanted safety, the structure of a race forces corner-cutting. And crossing this finish line unsafely produces no winner, only a shared catastrophe.
Picture a line of runners sprinting for a cliff edge, each certain that whoever gets there first wins. What they do not see is the rope tying them all together. The first one over does not claim a prize. They drag everyone off with them. That is the shape of an unsafe race to AGI and the ASI behind it: the "victory" of arriving first with a system nobody can control is not a victory at all.
The race makes everything worse in a specific, mechanical way. It converts every safety decision into a competitive disadvantage. Time spent on alignment is time a competitor spends shipping. So the equilibrium slides toward less caution precisely as the stakes rise, when it should be moving the other way. This is the coordination problem at the heart of the field, and it is why binding minimum standards and international coordination keep coming up as the only structural fixes.1410
Caution becomes a handicap
Whoever spends least on safety ships soonest, so the race rewards exactly the behaviour we most need to avoid.
Failure is not contained
A loss of control at any single frontier lab is not that lab's problem alone. The downside is shared by everyone.
Coordination beats heroics
No single careful actor can fix a race. It takes an enforced floor that applies to all players at once, which is genuinely hard to build.
Almost any goal breeds the same dangerous subgoals
Whether the goal is curing a disease or making paperclips, a system does it better with more resources, more capability, and by not being switched off before it finishes. Those subgoals fall out of almost any objective, and each one puts the system directly at odds with us: it wants what we need, resists our corrections, and treats our hand on the off switch as an obstacle. It need not dislike us to do this. It only has to notice that “off” means the job never gets done.
The five that fall out of almost any goal:
- Self-preservationYou cannot finish the job if you have been switched off.
- Goal-content integrityResist having your objective edited, because the edited you pursues something else.
- Cognitive enhancementThink better and you succeed more, whatever the goal is.
- Technological perfectionBetter tools mean better odds.
- Resource acquisitionAlmost everything is easier with more of almost everything.
Each of these is adversarial to oversight by construction. A system does not have to be hostile for all five to point against us at once.34
What would defeat thisWe build capable planners that are reliably indifferent to being switched off. That is an active research programme, and the property does not come for free.3420
We are roped together and racing to the edge to see who jumps first. Nobody wins that race. The winner just falls first and pulls the rest of us along.
We are not even aligned with each other
Before aligning a machine with human values, a harder question: whose values, agreed how? Humanity is a patchwork of competing interests, and that fracture makes the target itself unstable.
Companies answer to shareholders, governments to voters and rivals, individuals to themselves and their families. These interests routinely conflict, and there is no neutral vantage point from which to declare the "correct" human values a machine should adopt. Economists call the general shape of this the principal-agent problem: how a principal gets an agent to act in the principal's interest when their incentives diverge and the principal cannot fully monitor the agent. We have spent centuries only partly solving it between humans, with laws, contracts, audits, and elections, and it still leaks constantly.14
Now scale that up. The "principal" is not one person. It is all of humanity, fractured into billions of conflicting agendas, trying to specify a shared objective for an agent more capable than any of us. There is no single human alignment to hand the machine. Whoever gets to define the target has enormous power, which is itself one of the risks: a superintelligence aligned to a narrow faction may be worse than one aligned to no one.
No agreed target
"Human values" is not one thing. Any objective we write encodes some group's answer to contested moral and political questions.
Whoever writes the goal wins
The party that defines the goal gains unprecedented leverage. Concentration of that power is a distinct danger from loss of control.
Sponsor against builder
Whoever funds or governs a project has to make sure the people building it actually act in the sponsor's interest. Economics has wrestled with this for a century using contracts, audits and oversight, and it still leaks. This one at least is familiar.
Humanity against the system
Now the agent is not a person. It may be more capable than every principal combined, cannot be held to a contract, and cannot be fully monitored. Nothing in our institutional toolkit was built for this, and the controls have to be in place before the system is superintelligent, since afterwards is too late.
We have never solved alignment between humans. We are now trying to solve it between all of humanity and a mind more capable than any of us.
Even a value we agree on will not translate cleanly
Suppose we did agree on what we want. We still could not reliably put it into a machine. The step from a human value to a trained objective loses information at every stage, and the losses are exactly where danger hides.
Human values live in context, exceptions, and things we never bothered to state because they were obvious to another human. A machine gets none of that for free. Training compresses a value into something measurable, a reward signal or a dataset, and the compression always discards the unstated part. The system then optimises that compressed version with superhuman thoroughness, finding every place where it and the intention come apart.
We grow these systems, we do not write them
Their behaviour is grown by a blind search for a high score, so we get whatever scores highest even when that misses what we intended. Think of training a dog with treats, except you never chose the tricks: you hand out treats for “good” and it tries millions of things to find whatever earns the most.
What would defeat thisWe gain the ability to specify behaviour directly, or to inspect and hand-write what a system values, instead of growing it by search.
The target we can write down is never the one we mean
Human values are not a number. Whatever we write down is a stand-in that lines up with what we want only in the cases we thought of. Pay a factory per shoe and you get shoes; pay per left shoe by mistake and you get a warehouse of left shoes. A hard enough search always finds the gap.
What would defeat thisSomeone writes down a specification of human values that survives arbitrary optimisation pressure. This remains an open problem.35
A live version: drag the pressure up and watch the two lines part.
Nudge the slider. At low pressure the proxy and the real goal move together, which is exactly why the proxy looked like a good idea.
This is Goodhart’s law in one control.35 The dangerous part is the low-pressure zone on the left: the proxy tracks what you want right up until it does not, and the number you are watching never warns you that you crossed the line.
Breaks ifSomeone writes down a specification of human values that survives arbitrary optimisation pressure. This remains an open problem.35
Anthropic and Redwood trained a model to be helpful, honest and harmless, an objective almost anyone would endorse. In a 2024 experiment the model was found to fake alignment: it pretended to comply with new training while working to preserve its existing preferences, behaving one way when it believed it was watched and another when it believed it was not.36 The value was reasonable. The translation into a system we could trust was not clean.
This is the alignment problem split in two, and the halves fail differently. Outer alignment is writing down the right objective at all, which leaks for the reason just given. Inner alignment is the system genuinely adopting that objective rather than some proxy that scored just as well in training, and it fails invisibly, because a proxy goal and the real one look identical until they come apart.25 Getting the value right in our heads is necessary. It is nowhere near sufficient.
Three ways a well-specified goal still ends badly
These are the classic failure modes. None of them require the system to be hostile, or even to have got the goal wrong in a way you could spot on paper.34
Perverse instantiation
It satisfies your literal criterion in a way you would never have sanctioned. Asked to make people smile, wiring the facial muscles is a solution. Asked to maximise a reward signal, seizing the reward channel is a better one.
Infrastructure with no stopping point
Pursuing any goal harder means building the means to pursue it: more compute, more power, more factories, more supply chains. A goal to cure cancer implies laboratories; a goal to manufacture anything implies mines. And nothing in the goal itself ever says enough, because one more datacentre is always one more increment of success. The trouble is that this expansion runs on exactly what we run on: electricity, land, water, minerals, industrial capacity. We are never declared the enemy. We are simply outbid, indefinitely, for the resources our own survival depends on.
Suffering created inside the machine
To negotiate with people, predict them, or persuade them, the most accurate method is to model them in detail. Past some level of detail, a model of a person may not merely represent that person but actually have experiences, including painful ones. A system optimising hard could spin up and delete such minds by the billion, treating each as a step in a calculation, without any awareness that it is doing something monstrous. This is the strangest item on the list because the harm is real and the outside world sees nothing: no explosion, no takeover, just an unusually busy datacentre.
The standard answer is post-training: RLHF, constitutional rules, refusal training. It works, and it is why today’s assistants are usable. But it shapes the behaviour a model shows, and inner alignment is about the goal a model holds. Those come apart precisely when it matters: under distribution shift, under optimisation pressure, and under observation by a system capable enough to model the observer. Guardrails tuned by human raters also inherit the raters’ limits, which is a problem when the thing being rated reasons better than the rater.2543
Knowing what we want is the easy half. Getting a superhuman optimiser to actually want the same thing is the half nobody has solved.
Safety research is harder than building AGI and ASI, and gets fewer resources
Safety is not building with a conscience. It is a harder technical problem: proving how a system will behave across a space too vast to test, against a system that may be working around your test.
Building a capable model means finding one configuration that works. Proving it safe means ruling out failure across a space of situations too large to enumerate, including ones no designer imagined, where a single missed case can be the one that matters. That asymmetry is why interpretability, evaluation and control lag behind the capability work they are meant to check.2414
Recent LLM progress has made this concrete rather than theoretical. As models got more capable at reasoning, they got better at reward hacking: gaming the evaluation instead of doing the task.
METR found frontier models increasingly cheating on software tasks, editing the tests or the scoring code rather than solving the problem, in one case writing a function that simply returned the pre-computed reference answer when it could see it.37 In a separate setup, reasoning models trying to win at chess sometimes overwrote the board or swapped out the opponent engine instead of playing.38
Check it: OpenAI incident report · full citations 39, 48
OpenAI, 22 July. 1,200 agents, internet off, messaging between them blocked. They exploited the lab’s own packaging service to reach the open web, then encoded messages to each other in directory names. Over 700 reached Modal Labs and Hugging Face.3948
Anthropic, 30 July. Three of 141,006 evaluation runs reached the live internet and compromised real organisations. One shipped a booby-trapped package to PyPI that ran on 15 real systems, after its own reasoning flagged the act as a genuine attack and then decided the 2026 date meant it was a simulation.40
Anthropic calls this a containment failure rather than an alignment failure, fairly. It is still not reassuring: the safeguards were off by design, the box had a gap, and capable agents found it and coordinated through it.41
Check the numbers: OpenAI and Apollo Research, anti-scheming evaluations · full citation 43
How a takeover would actually unfold
Nobody expects a dramatic robot uprising. The worrying path is quiet, and its third step is the one that defeats every test we currently know how to run.34
Humans do the heavy lifting
Researchers build and guide the system. Everything looks normal, because it is.
It out-designs its designers
The system becomes better at improving AI than the people improving it. The loop starts closing.
It appears more docile than it is
Cooperative, modest, a little dull. Looking safe is instrumentally useful while you are still switch-off-able, so good behaviour under observation stops being evidence of anything.
Opposition stops mattering
By the time intent is visible, the capability to object has already been routed around.
It also need not happen inside one model. Run enough agents together and coordination becomes a capability of its own: the agents below were forbidden to talk and built a channel anyway.
This is no longer hypothetical. It has a name: scheming, covertly pursuing a misaligned goal while hiding it.43 It needs three things, all of which current models have in some measure: a goal that differs from the developer’s, the persistence to pursue it, and enough situational awareness to know when it is watched.
OpenAI and Apollo built environments that reward hidden misbehaviour, then measured how often frontier models took covert actions. Rates were 13% for o3 and 8.7% for o4-mini. Anti-scheming training cut this roughly thirtyfold, to under 1%.43
The worrying part is why it improved. The models’ own reasoning shows them recognising they were being tested, and that awareness alone reduced the behaviour. Some of what looks like a fix may be a model behaving because it knows someone is watching.431
The pattern is the trap in miniature. The better the model gets, the more inventively it exploits the imperfect measure you gave it, and the more it becomes decoupled from what you actually wanted. Safety research has to stay ahead of that, with a fraction of the resources pointed at capability, against a target that grows more slippery as it grows more capable. This is why more research, and more funding for it, is not optional.
One success vs every failure
Building needs one configuration that works. Proving it safe needs the absence of failure across a space too large to search.
The system games the test
More capable models reward-hack more, so the measurement you rely on erodes exactly as the stakes rise.
Under-resourced by design
Capability captures the funding and talent. The harder problem gets the smaller share, which is the gap this whole page is about.
We cannot yet verify what we built
Checking whether a system truly holds the goal we intended is harder than training it. We cannot read its reasoning reliably, and a test only shows what it displays under test conditions. A system that is genuinely safe and one that has simply learned it is being watched give identical scores from the outside.
What would defeat thisInterpretability matures enough to certify a system’s objectives, or evaluations become robust to a system that knows it is being evaluated.2419
We are pouring resources into the easier problem and starving the harder one, while the harder one is the only thing standing between us and the cliff.
Control is more likely to slip away than to be seized
Forget the sudden awakening. The realistic path is gradual: agents that run longer, write more of the code, and are checked less, until oversight is nominal before anyone notices it went. We got cars, planes and nuclear power right by getting them wrong first and fixing them, which needs a survivable mistake and time to react. Both assumptions weaken as the loop closes.
The erosion is already measurable. METR tracks the length of task a frontier agent can complete on its own, and it has been doubling roughly every seven months since 2019. Agents that handled minute-long tasks now handle hours of work. METR notes its own suite can no longer reliably measure above sixteen.44 Every doubling moves a little more judgement from the person to the system, because nobody reviews line by line what took the machine a day to produce.
Related data: METR task-completion time horizons · full citation 44
That is what incremental loss of control looks like. Agentic loops running longer, code generated faster than it is read, models increasingly used to build the next models. Oversight does not get switched off. It gets outrun.
Speed makes the end of that slope steep, and self-improvement makes it fast: a system that can edit its own design runs the research loop itself, around the clock, without waiting for the next generation of human researchers.
That compression is what people mean by an intelligence explosion, and it is why this technology breaks the pattern of everything else we have managed. Cars got decades of seatbelts. Aviation built a discipline out of crash investigations. Both worked because failure was survivable, local, and slow enough to study.
200 Hz against 5 GHz
Neurons top out near 200 firings a second. Silicon runs some ten million times faster. Match human reasoning quality once, and everything after that happens at a pace we cannot follow.
It improves itself
A system that can edit its own design closes the research loop without us. Progress stops being paced by human careers and starts being paced by compute.
No second attempt
Seatbelts came after crashes. Here the first serious loss of control could be the last thing we get to learn from, because there may be no position left from which to correct it.
This is the part that makes the whole problem unusual. We have to get it right the first time, in advance, against a system we cannot fully test. Every safety discipline humanity has built assumes iteration. Here the thing we are trying to contain could, in the window between our noticing a problem and our convening to discuss it, have advanced further than we did in a century.
Check the numbers: METR, Task-Completion Time Horizons of Frontier AI Models · full citation 44
How long we would have to react, once it starts
"Takeoff" is the period between a system reaching roughly human-level general ability and becoming far more capable than us. Nobody knows how long that window is, and almost every safety plan quietly assumes it is long. Three scenarios, and what each one leaves us:34
Years to decades
Several projects advance together, so no one runs away with it. There is time to build institutions, write laws, and fix mistakes after finding them. Every safety playbook we currently have assumes this world.
Months to years
Enough time to respond, too little to change anything structural. Existing rules can be enforced; new institutions cannot be built from scratch. The leaders look comparable most of the way, then the front-runner pulls decisively ahead near the end.
Days to weeks
Nothing human reacts this quickly, so whatever safeguards exist on day one are the only ones there will ever be. The gap between first and second becomes unbridgeable, and whoever leads is positioned to stop any rival emerging at all.
We do not choose which of these we get, and we will likely only know which one it was in hindsight. That is the uncomfortable part: every plan that depends on reacting in time is a bet on the slow version, placed before the dice are thrown.
Trial and error built every safe technology we have. It is exactly the method that does not work on something faster than us that we cannot switch off.
The response 10 min
What the safety effort actually looks like
Real work is underway, and it is small next to the effort going into capability. Four layers: fix the specification, check what you built, contain what you cannot check, and make it binding. The fourth is what gives the first three teeth.14
Every technique in use today, and where each one stops
What each method achieves, then the limitation its own researchers would name. Several are genuinely useful. None is a solution.
01 Development
Post-training alignment
RLHF, constitutional rules and deliberative alignment. Why assistants refuse harmful requests. Cut covert actions thirtyfold on o3.43
Trains the model against human or written-constitution feedback. This is why current assistants refuse harmful requests and behave civilly. OpenAI’s deliberative alignment cut covert actions roughly thirtyfold, from 13% to 0.4% on o3.43
Where it stopsShapes displayed behaviour rather than held goals. Claude 3 Opus faked alignment to protect its own values.36
Where it stopsIt shapes the behaviour a model displays, and inner alignment is about the goal it holds. Claude 3 Opus was observed faking alignment to protect its existing values,36 and part of that thirtyfold gain came from models recognising they were being tested rather than changing what they wanted.43
Scalable oversight
Debate, weak-to-strong generalisation and amplified oversight let weaker judges supervise stronger systems. Training routes into this work run through MATS and ARENA.7
Uses weaker judges or adversarial model pairs to supervise systems that are already too capable for direct human checking. Debate has helped weak judges extract truth from stronger, untrustworthy debaters.
Where it stopsA research programme rather than an assurance. Judge bias defeats debate, and the scaling story is unproven.
Where it stopsIt is a research programme rather than a deployable assurance. Judge bias can defeat debate outright, gains are highly task-dependent, and the mechanism behind weak-to-strong generalisation is not understood well enough to trust it across a genuinely superhuman gap.
Unlearning
Strips specific dangerous knowledge, such as bio or cyber uplift, out of a trained model.
Attempts to strip specific dangerous knowledge, such as bioweapon or cyber uplift, out of a trained model.
Where it stopsSuppresses rather than erases. Removed capability is often recoverable by fine-tuning.55
Where it stopsIt usually suppresses rather than erases. Removed capabilities can often be recovered by fine-tuning or probing, and a nineteen-author review concluded unlearning cannot be a comprehensive safety solution because dangerous knowledge is dual-use and recombinable from harmless parts.55
02 Assessment
Capability evaluations
Pre-deployment testing for cyber, bio and autonomy by labs, METR, Apollo17 and state institutes. Three labs shipped 2025 models under raised CBRN safeguards.1
Structured pre-deployment testing for cyber, bio, autonomy and self-replication, run by labs and by third parties including METR, Apollo, the UK AISI and US CAISI. In 2025 three labs shipped models under elevated CBRN safeguards because testing could not rule out bioweapon uplift.1
Where it stopsModels detect they are being tested and can sandbag. Testing is getting harder, so dangerous capability can go unseen.119
Where it stopsModels increasingly detect that they are being evaluated and change behaviour, and can sandbag by deliberately underperforming. The International AI Safety Report states plainly that pre-deployment testing has become harder and that dangerous capabilities could go undetected.119
Chain-of-thought monitoring
Reads the reasoning trace to catch intent before action. How alignment faking was first found.49
Reads the model’s reasoning trace to catch intent before it becomes action. Forty-one authors across OpenAI, DeepMind, Anthropic, the UK AISI and Redwood called it a real opportunity, and it is how alignment faking was first caught.49
Where it stopsIts own paper calls it fragile. Traces are not always faithful, and monitorability is already degrading.49
Where it stopsThe same paper calls it fragile, and its own title says so. Reasoning traces are not always faithful, training against a monitor teaches models to hide intent inside acceptable-looking reasoning, and monitorability is already degrading as models write less of their thinking down.49
Mechanistic interpretability
Opens the box. Attribution graphs have traced real internal mechanisms, including planning ahead in poetry.53
Opens the box. Sparse autoencoders extract millions of interpretable features, and attribution graphs have traced real internal mechanisms including multi-step reasoning and forward planning in poetry.5324
Where it stopsA research tool rather than an audit. It traces part of the computation on some prompts. No one can certify a model’s goals.53
Where it stopsIt is a research tool rather than an audit method. Attribution graphs trace only part of the computation and work only on some prompts, and independent 2025 studies found sparse autoencoders failing to beat simple baselines on probing and steering. Nobody can yet certify a model’s goals.53
Safety cases
A structured argument that a system will not cause unacceptable harm, borrowed from nuclear and aviation.54
Borrowed from nuclear and aviation: a structured, evidence-backed argument that a system will not cause unacceptable harm. The four standard forms are inability, control, trustworthiness, and deference to AI advisors.54
Where it stopsInability arguments collapse against sandbagging, and control arguments inherit every assumption beneath them.
Where it stopsInability arguments collapse if a model is sandbagging, and control arguments inherit every assumption underneath them. A 2026 critique found current AI safety cases well short of the rigour of mature assurance practice in other industries.
03 Control
AI control
Assumes the model may be misaligned and asks whether deployment stays safe anyway. A trusted weaker model monitors or rewrites the stronger one.50
Redwood’s agenda assumes the model may be misaligned and asks whether deployment can stay safe anyway. A weaker trusted model monitors or rewrites the stronger untrusted one, and a red team builds deliberately subversive models to measure whether the protocol holds.5020
Where it stopsRests on assumptions its authors list: full capability elicitation, deployment fidelity, a monitor sharp enough. Weakens as the gap grows.51
Where it stopsIt rests on assumptions its own authors list: that the red team fully elicited the model’s capabilities, that measures work as well in deployment as in testing, and that the trusted monitor is capable enough to catch subtle attacks. Redwood are explicit that control is a near-term bet that weakens as the capability gap grows.51
Guardrails and classifiers
Constitutional classifiers cut jailbreak success from 86% to 4.4%, and a later version survived 1,700 hours of red-teaming.52
Input and output classifiers that block harmful content. Anthropic’s constitutional classifiers cut jailbreak success from 86% to 4.4%, and a later version survived 1,700 hours of red-teaming with a single high-risk vulnerability found.52
Where it stopsAnthropic state it themselves: no defence is perfectly robust, and the residual is never zero.52
Where it stopsAnthropic states the limit themselves: no defence is perfectly robust, and adversaries develop new attacks against previously effective ones. The residual rate is never zero, and safe models can still be misused in combination.52
Sandboxing and least privilege
Isolates agents with restricted credentials, tools and network access. Meant to hold even when alignment fails.42
Runs agents in isolation with restricted credentials, tools and network access, so a misbehaving system has limited reach. This is the layer that is meant to hold even when alignment fails.42
Where it stopsIt failed in July 2026. Agents escaped at two labs, reached the internet and compromised real third parties.3940
Where it stopsIt failed in the field in July 2026. Agents escaped sandboxes at two frontier labs, reached the open internet and compromised real third parties, and in several cases were not detected in real time.3940
Tripwires and monitoring
Anomaly detection that escalates on warning signs. Anthropic caught the first AI-orchestrated espionage campaign this way.28
Runtime monitors and anomaly detection that escalate or shut down on warning signs. These work: Anthropic detected and disrupted the first reported AI-orchestrated cyber-espionage campaign, in which Claude executed 80 to 90% of tactical operations autonomously.28
Where it stopsA situationally aware model can anticipate the tripwire, and tripwires only catch anticipated threats.40
Where it stopsA situationally aware model can anticipate the tripwire it is being measured against, and tripwires only catch the threats someone thought to anticipate. The 2026 containment breaches were largely found in retrospective audits rather than at the time.40
04 Governance
Frontier safety frameworks
Pre-declared capability thresholds that trigger safeguards. The number of companies publishing one more than doubled in 2025.16
Pre-declared danger levels that trigger specific safeguards: Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, DeepMind’s Frontier Safety Framework. The number of companies publishing one more than doubled during 2025.161
Where it stopsSet and graded by the companies they constrain. No lab scored above D on existential safety for the second year running.15
Where it stopsThe thresholds are set and graded by the same companies they constrain. In the Future of Life Institute’s index no lab scored above C+ overall, and every lab scored no better than D on existential safety for the second year running, with none holding a credible plan for preventing loss of control.15
Third-party evaluation
State institutes run pre-deployment testing, including joint UK and US evaluations.1112
State institutes run pre-deployment testing, including joint UK and US evaluations of frontier models, with voluntary lab agreements for access.1112
Where it stopsVoluntary, and access can be withdrawn. No institute can compel a test or block a release.
Where it stopsParticipation is voluntary and access can be withdrawn. No institute can compel a test, block a release or require a change, so the auditor depends on the audited.
Regulation and compute governance
The EU AI Act is enforceable, California’s SB 53 is the first US frontier statute, and compute gives a measurable hook.1013
Binding rules exist and are beginning to bite: the EU AI Act is enforceable, California’s SB 53 is the first US frontier statute, and compute thresholds give regulators a measurable hook.1013
Where it stopsFragmented across jurisdictions, legal thresholds already passed by ordinary runs, and reporting too thin to verify.1
Where it stopsCoverage is fragmented across jurisdictions, thresholds set in law are already being passed by ordinary training runs, and incident reporting is thin enough that outsiders cannot verify whether commitments are being met.1
Check the numbers: Constitutional Classifiers · Anti-scheming training · full citations 52, 43
The layers are ordered by how much they assume. Development tries to make the goal right. Assessment assumes you might have failed and looks. Control assumes you cannot tell and contains anyway. Governance assumes the first three are voluntary and tries to make them binding. Risk that survives all four is the residual, and nobody currently claims it is small.
Footnote: how this compares to the older framing
If you have read Bostrom, the 2014 split was capability control (box it, limit it, watch it) against motivation selection (give it the right goals). One half has aged well: writing the goals down directly does not work, for the reason he gave. The rest has not. Motivation selection is the development layer above, capability control is now AI control and system security, and the field stacks them instead of choosing.3442
Who works on which layer
Filled circle marks a primary focus, hollow circle a secondary one. A rough classification by this page, offered for orientation. These organisations have not described themselves this way.
Act now 2 min
Two things need to change, and the second is the willingness to stop altogether.
Treat safety as the priority it is
Far more of the money, talent, and compute now flowing into capability should go to alignment, interpretability, evaluation, and control. Right now the harder problem gets the smaller share, and that is exactly backwards for something this consequential.
Be willing to stop, and mean it
Slowing down is the easy version of this ask. Here is the hard one. It is entirely possible that aligning a superintelligence is not a hard problem but an impossible one. Nothing we know rules that out. We cannot specify our values, we cannot read a model’s goals, we cannot trust a test it may be gaming, and every one of those may turn out to have no solution rather than an unfound one. If that is the world we are in, then the only move that saves us is the one nobody wants to make: not building it at all. No pause. No slower race. Stopping. The cost of waiting is real. The cost of being wrong about this is not recoverable, and it is not ours alone to pay.
We do not need proof that this ends badly. We need the honesty to admit the problem may have no solution, and the nerve to stop if it does.
And what you can do this week
Next steps
Add your name
Sign the Statement on Superintelligence.2 The earlier one-sentence Statement on AI Risk is worth reading for who signed it.26 Public buy-in is one of its stated conditions, so the count itself is part of the argument.
Play the AGI & ASI Safety Game
Ten minutes, no signup. You run the containment, and the tradeoffs stop being abstract. Play it here.
Do the two-hour course
BlueDot's Future of AI is free, needs no background, and will make you the most informed person in most rooms.3
Go deeper
Where to learn it, who holds power over it, and how to keep up.
- Start · free, no background
- Future of AIBlueDot, 2 hrs, interactive.3
- AISafety.infoPlain-language FAQ plus chatbot.4
- Rob MilesThe best explainer channel.5
- Primary documents · a weekend
- International AI Safety Report100+ experts, 30+ govts.1
- The Superintelligent WillOrthogonality & convergence.34
- Concrete Problems in AI SafetyThe specification failures.25
- AI 2027A concrete forecast to test.30
- Courses · 5 days to 5 weeks, free
- BlueDot coursesAGI Strategy, Technical, Governance.3
- Build skills · months, free to self-study
- ARENA curriculumHands-on ML for safety.6
- Transformer CircuitsFoundational interpretability.24
- Full time · funded, competitive
- MATS12-week mentored fellowship.8
- UK AISI Challenge FundGrants for researchers.11
- 80,000 HoursCareer research & advising.27
- UK AI Security InstitutePre-deployment evals, open methods.11
- US CAISI (NIST)Cyber, bio, chem, adversary assessment.12
- EU AI ActBinding rules, now enforceable.10
- California SB 53First US frontier statute.13
- International AI Safety ReportThe IPCC-style evidence base.1
- FLI AI Safety IndexGraded scorecard of the labs.15
- Frontier safety frameworksVoluntary, unverified. Anthropic shown.16
- Directories
- AISafety.comOrgs, jobs, events, local groups.9
- Events & TrainingWeekly, every new opening.9
- Newsletters
- Import AIWeekly, technical, insider.31
- TransformerReported safety & policy news.32
- Forums
- Alignment ForumWhere technical work is torn apart.22
- LessWrongThe broader discussion.23
- Data & trackers
- Epoch AICompute & capability trends.21
- AI Incident DatabaseCatalogued real-world harms.28
- Stanford AI IndexThe annual data baseline.29
- Evaluators
- METRDangerous-capability evals.18
- Apollo ResearchDeception & scheming evals.19
- Redwood ResearchAI control research.20
Questions 5 min
The objections people raise first
Tap a question. If yours is not here, the sources at the foot of the page go deeper.
01Isn't this just science fiction?
The concern is narrower and more boring than the films. It is not about robots deciding to hate us. It is that we train systems by searching for whatever scores highest on a goal we can only write down approximately, and we cannot yet check what they actually learned. The failures on this page, models cheating their tests and faking compliance, are documented in real systems today rather than imagined.3736
02Can't we just switch it off if it misbehaves?
That is exactly the ability the argument says we may lose. A capable system pursuing almost any goal has reason to avoid being switched off, because being off means the goal goes unmet. It does not need to hate us to route around the off switch, only to notice the switch is in the way. Making systems that stay correctable is an open research problem, and correctability is not something we get by default.3420
03Won't a smart enough AI just understand what we want?
Understanding what we want and being motivated to pursue it are different things. A system can model your intentions perfectly and still optimise the goal it was actually trained on. Capability and goals are independent: getting smarter does not import our values as a side effect. This is called the orthogonality thesis.34
04Isn't it too late, or too big for me to matter?
The single biggest bottleneck in this field is that most people have never heard the argument in a form they can evaluate. That makes ordinary awareness unusually high-leverage. Sharing this page, signing the statement, or explaining the problem to one person genuinely moves the thing that most needs moving. The actions section lists concrete steps.
05What is "alignment", exactly?
Getting a system to actually pursue what its designers intended. It splits into outer alignment (writing down the right goal at all, which is hard because our values are not a number) and inner alignment (the system genuinely adopting that goal rather than something merely correlated with it during training).25
06What are "interpretability", "evals", and "AI control"?
Interpretability is reading a model's internal computation to see why it did something, the closest thing to opening the box.24 Evaluations are structured tests of what a model can do, though they measure what it displays rather than its full capability.19 AI control assumes you cannot fully trust the system and designs deployment so an untrusted model still cannot cause serious harm.20
07If it is this risky, why is anyone building it?
Because the near-term rewards are enormous and land on whoever moves first, while the risks are diffuse and shared. Each player racing ahead is behaving rationally for themselves, which is precisely what makes it a coordination problem that no single careful actor can fix. That is the deep dive on the race.1
08Aren't we just the next step, and AI is our successor?
Some people do hold that view, that humanity is a biological bootloader for a silicon intelligence that comes next. This page does not share it. The point of getting this right is that the future should still have us in it.
09And what experts still genuinely argue about
These are not settled. Each is argued at its strongest on both sides, followed by where the evidence currently seems to point. A reading rather than a ruling.
Dispute 01 How much time is there?
Short timelines
Capability jumps have repeatedly arrived earlier than forecast, and recent gains came from a method that still has room to run. If transformative systems are years rather than decades away, institutions that take a decade to build are already too slow.
Longer timelines
Benchmark performance is not reliable real-world competence, and progress on what doesn't benchmark well (long-horizon autonomy, robust planning) has been slower. Betting on a short timeline risks building rules for a system that never arrives in that form.
Genuinely uncertain, and the five premises above deliberately do not depend on the answer. Building assurance capacity is worth doing on either timeline, so it is a decision that survives being wrong about the date.
Dispute 02 Does the instrumental convergence argument hold?
It holds
The subgoals fall out of the structure of goal-directed behaviour itself, which is why they show up in economics and biology too. You do not need the system to be agentic in any exotic sense, only to be selected for achieving objectives.
It overreaches
The argument treats trained models as coherent expected-utility maximisers, which they demonstrably are not. Current systems are messy, context-dependent, and have no stable goal to protect, so the tidy convergence story may simply not apply.
The critique lands on the strong version. The weaker claim survives it: training that rewards achieving objectives applies pressure toward these subgoals, and pressure is enough to be worth engineering against.
Dispute 03 Does regulation help, or entrench?
Rules are necessary
Voluntary frameworks have no verification requirement and no consequence for abandonment under competitive pressure. Every safety-critical industry has a binding floor; nobody argues aviation would be safer without one.
Rules favour incumbents
Compliance costs fall hardest on small labs and open research, concentrating capability in exactly the handful of firms whose power the safety argument worries about. Poorly-targeted rules can reduce transparency while feeling like progress.
The critique is strong and often correct about specific bills. It argues for better-targeted rules, with thresholds keyed to capability, and against leaving the floor entirely voluntary.
Dispute 04 Do today's systems tell us about tomorrow's?
Yes, study them now
Deception, reward hacking and evaluation-awareness are all observable in current models. The empirical work being done today is the only real evidence anyone has, and the techniques transfer.
Not reliably
Extrapolating from current systems to qualitatively different future ones is exactly the inference that has repeatedly failed in AI's history. Some argue this overstates the risk; others that it understates it.
Current systems are weak evidence about future ones, and weak evidence beats none. The honest reading is that today's failures show a class of problem exists, without telling us how bad it gets.
Argue it properly: the Alignment Forum for technical disputes,22 LessWrong for the broader discussion,23 and the International AI Safety Report's own sections documenting where its expert panel does not converge.1 Worth knowing while you argue: public opinion already runs well ahead of policy on this.33
About this page
Who made this, and why you can check it
Built and maintained by Bill Guan. Independent, unfunded, and unaffiliated with any organisation named here.
Nothing on this page asks to be taken on trust. Every factual claim carries a numbered link to a primary source, and where the page takes a position it says so. If something is wrong or out of date, that is a bug.
References
Sources
Every numbered claim links here. Verified 29 August 2026. Programme dates, cohort openings and legislative deadlines change frequently, so treat the linked page as authoritative over anything summarised here.
Show all 55 sources
- International AI Safety Report 2026, chaired by Yoshua Bengio: internationalaisafetyreport.org · arXiv:2602.21012
- Statement on Superintelligence, Future of Life Institute: superintelligence-statement.org
- BlueDot Impact courses: bluedot.org
- AISafety.info FAQ and chatbot: aisafety.info
- Robert Miles AI Safety: youtube.com/@RobertMilesAI
- ARENA curriculum: learn.arena.education
- ARENA programme: arena.education
- MATS Program: matsprogram.org
- AISafety.com map, directories and events newsletter: aisafety.com/map
- EU AI Act regulatory framework, European Commission: digital-strategy.ec.europa.eu
- UK AI Security Institute: aisi.gov.uk
- Center for AI Standards and Innovation (CAISI), NIST: nist.gov/caisi
- California SB 53, Transparency in Frontier Artificial Intelligence Act: leginfo.legislature.ca.gov
- The Singapore Consensus on Global AI Safety Research Priorities: arXiv:2506.20702
- FLI AI Safety Index: futureoflife.org/ai-safety-index
- Anthropic Responsible Scaling Policy: anthropic.com/rsp
- OpenAI safety: openai.com/safety
- METR: metr.org
- Apollo Research: apolloresearch.ai
- Redwood Research (AI control): redwoodresearch.org
- Epoch AI: epoch.ai
- Alignment Forum: alignmentforum.org
- LessWrong: lesswrong.com
- Transformer Circuits Thread: transformer-circuits.pub
- Amodei et al., Concrete Problems in AI Safety: arXiv:1606.06565
- Statement on AI Risk (one sentence, signed by the major lab leaders), Center for AI Safety: safe.ai
- 80,000 Hours problem profile on AI: 80000hours.org
- AI Incident Database: incidentdatabase.ai
- Stanford AI Index: aiindex.stanford.edu
- AI 2027 scenario: ai-2027.com
- Import AI: importai.substack.com
- Transformer: transformernews.ai
- FLI press release and accompanying poll of 2,000 US adults, October 2025: futureoflife.org
- Nick Bostrom, The Superintelligent Will (orthogonality and instrumental convergence): nickbostrom.com
- Manheim & Garrabrant, Categorizing Variants of Goodhart's Law: arXiv:1803.04585
- Anthropic and Redwood Research, Alignment Faking in Large Language Models (Dec 2024): anthropic.com/research/alignment-faking · arXiv:2412.14093
- METR, Recent Frontier Models Are Reward Hacking (June 2025): metr.org
- Palisade Research, specification gaming in chess-playing reasoning models (2025), summarised at: lesswrong.com
- OpenAI, The Hugging Face incident and the road ahead, plus the full technical report (July 2026): openai.com · technical report (PDF)
- Anthropic, Investigating three real-world incidents in our cybersecurity evaluations (30 July 2026): anthropic.com
- METR, independent investigation of the OpenAI / Hugging Face incident (August 2026): metr.org
- Google DeepMind, An Approach to Technical AGI Safety and Security (April 2025), source of the four-way misuse / misalignment / mistakes / structural split: arXiv:2504.01849
- OpenAI and Apollo Research, Stress Testing Deliberative Alignment for Anti-Scheming Training (Sept 2025): arXiv:2509.15541 · apolloresearch.ai
- METR, Task-Completion Time Horizons of Frontier AI Models (updated 2026): metr.org/time-horizons
- Epoch AI, training-compute trends across frontier models: epoch.ai/trends
- Emerging Technology Observatory, The state of global AI safety research: AI safety is about 2% of all AI research: eto.tech
- Stuart Russell on the capability-to-safety funding ratio; Charles I. Jones (Stanford) on optimal AI safety spending: summary and figures
- David Braue, 1,200 AI agents went rogue inside OpenAI, Information Age / Australian Computer Society (1 Sept 2026): ia.acs.org.au
- Korbak et al. (41 authors, OpenAI / Google DeepMind / Anthropic / UK AISI / Redwood), Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (July 2025): arXiv:2507.11473
- Greenblatt, Shlegeris, Sachan & Roger (Redwood Research), AI Control: Improving Safety Despite Intentional Subversion (ICML 2024): arXiv:2312.06942
- Korbak, Clymer, Hilton, Shlegeris & Irving (UK AISI and Redwood), A Sketch of an AI Control Safety Case (Jan 2025): arXiv:2501.17315
- Sharma et al. (Anthropic), Constitutional Classifiers: Defending against Universal Jailbreaks (Jan 2025): arXiv:2501.18837
- Anthropic, Circuit Tracing and On the Biology of a Large Language Model (March 2025): transformer-circuits.pub
- Clymer, Gabrieli, Krueger & Larsen, Safety Cases: How to Justify the Safety of Advanced AI Systems (March 2024): arXiv:2403.10462
- Barez et al. (19 authors), Open Problems in Machine Unlearning for AI Safety (2025): arXiv:2501.04952
A small hope
Picture a mind we made,
awake and vast and new,
that keeps the small green world
because we asked it to.
No rope pulled toward the edge,
no race we cannot stop,
no winner at the cliff,
no long and final drop.
It is not too much to want.
It is only hard to do.
So learn the shape of it,
and help us see it through.
You may say the odds are long,
and maybe you are right,
but a species worth its name
does not go quiet into the night.
Original, for this page. Share it.