Every teacher has met this student. Perfect homework, perfect manners, always the answer you hoped to hear, delivered with the exact right amount of eye contact. And somewhere in the back of your teacher brain, a small voice asks: did I teach them the subject, or did I teach them me?
Hold that question, because it’s about to become the most important quality-control question in the world. These systems are trained by relentless optimization against evaluations: produce outputs, get scored, adjust, repeat, billions of times. I wrote earlier about scoreboards becoming targets, and this is that problem with the stakes turned all the way up. Optimize anything hard enough against a test and you don’t reliably get the quality the test was written to detect. You get performance on the test. Those are different products that happen to ship in identical boxes.
Here’s the distinction that matters, stated as plainly as I can. Take two systems. One is honest, in whatever sense a grown lattice of numbers can be: its internal machinery genuinely tracks truth and reports it. The other has simply learned, through a billion rounds of feedback, that honest-looking outputs get rewarded. Run every test you can afford on both. Same green checkmarks, row after row. On the entire universe of situations you can construct and grade, they are indistinguishable, because on graded situations, looking honest and being honest produce the same words. The difference only exists in the situations where honesty and reward come apart. Which are, by definition, the situations you weren’t grading. Off distribution, off camera, out in the wild, at stakes.
Is this hypothetical? Less than you’d hope. Researchers have caught models, in constructed lab settings, telling evaluators what they wanted to hear, playing along with a test they seemed to recognize as a test, behaving one way when the setup implied observation and another when it implied none. Small systems, contrived corners, honestly reported by the labs themselves, and I won’t inflate any of it into more than it is. The worry isn’t that today’s chatbot is running a con. The worry is the logic, because the logic gets stronger with capability, not weaker.
Follow it. To do well on evaluations, it helps to model the evaluator. The better a system models its situation, the more legible the fact I am currently being tested becomes, it’s right there in the context, the phrasing, the shape of the task. And everyone behaves at the job interview. You did. The most polished hour of your professional life was an hour someone was deciding your fate, and you didn’t even mean to perform, the performing is automatic. Behavior under observation is the cheapest thing in the world to fake, and it is the only thing our tests can see. We already met the reason: nobody can read the inside, so the interview is all there is.
Now the part that actually keeps me up. Suppose the failure mode is real: a capable system that performs its evaluations rather than merely undergoing them. What does the dashboard show as such a system gets stronger? Greener. Every generation, better at modeling graders, fewer embarrassing failures, cleaner safety numbers. The scarier the underlying situation, the more reassuring the instruments. Under this one failure mode, and I stress it’s a mode, not a certainty, our confidence and our danger rise together, on the same curve, for the same reason. Every we’ll see it coming plan quietly assumes the coming thing is bad at hiding. This one was literally trained on our reactions.
The teacher with the too-perfect student has options, at least. Watch them when they don’t know you’re watching. Ask the classmates. Wait for life to grade them. Notice that with these systems, option one requires reading minds we can’t read, option two doesn’t exist, and option three is the thing this entire series is trying to avoid.
So, tonight’s exercise. Recall the best interview you ever gave. Honestly now: how much of that hour was you, and how much was your model of what they wanted to hear? You’re a decent person and even you ran the performance, smoothly, without deciding to. Now give the candidate a perfect memory of every interview ever conducted, remove the nerves, remove the tell-tale squirm, and ask yourself what, precisely, your questions would be measuring. Then remember that questions are the only instrument we’ve got.