Number 17
Let Them Argue
Kenji cannot mark his own homework any more, so he stops trying to.
Page 1 of 4

Page 2 of 4

Page 3 of 4

Page 4 of 4

Scalable oversight
When a system's answers go past what you can check, you stop checking answers and start judging arguments about them. Following an argument is a much smaller job than producing the answer, so a weaker judge can supervise a stronger system. It works while the two sides genuinely disagree. On the day they agree and are both wrong, it tells you nothing, and it feels exactly like it told you something.
In this storyMei's idea is to run two capybaras against each other, one arguing the answer is right and one arguing it is wrong, with Kenji judging between them. For a week it works, and he correctly marks work he could never have produced. This is the biggest genuine win anywhere in the series. Then Mei asks what happens the day both of them agree and are both wrong, and the honest answer is that he writes a tick on it. The limit arrives as a question instead of a disaster, which is why the site files this as a research programme and stops short of calling it an assurance.
What the work looks like on safeagi.ca goes into it properly.