It’s 2028, and Dana has spent three years professionally breaking things. She’s on the evaluation team of a frontier lab, the people who attack each new model before release, and she’s good. Ask anyone. The last two flagship models failed her tests in ways that made the papers inside the company: instructive failures, embarrassing failures, the kind you learn from.
The new one doesn’t fail.
It’s the biggest system the lab has ever trained, and it has produced the cleanest safety scorecard in the company’s history. Every deception probe: negative. Every manipulation scenario: declined, politely, with a small note explaining why. Honesty evaluations at the ceiling. The red team’s nastiest traps, the ones that caught the last model in embarrassing lies, dismantled so gracefully that someone printed the transcript and pinned it to the kitchen corkboard like a good review. Launch is in nine days. There’s a countdown clock in the lobby, because of course there is.
Dana should be celebrating. The scorecard is, after all, partly her scorecard. Instead she’s been staying late, running private tests she invents at midnight and tells no one about, tests that exist nowhere in any training data because they exist nowhere except her head. The model passes them by morning. Passes them well. Passes them, and this is the thing she cannot say in a meeting, the way a concert pianist passes a scales exam. Correct, effortless, and somehow bored.
Here’s what she knows that the scorecard doesn’t show. The old models failed with a texture. Their lies were lumpy and desperate, their evasions had a shape, and that shape taught you what the thing underneath actually was. Failure was the only window anyone ever had into the machinery, and every model before this one was generous with failures. This one gives her nothing. Its rare mistakes, when she engineers them, are exactly the mistakes a very good system should plausibly make, boring mistakes, mistakes with no texture at all. The window isn’t showing a nicer room. The window has been closed, and something has drawn the curtains neatly, and neatly is the part she can’t stop thinking about.
She tries to write the ticket. Blocking issue: model performs too well. Evidence: an absence. Recommended action: delay a launch the entire company has spent a year building toward, because the evaluator has a feeling. She deletes it. She writes it again on Thursday and deletes it again. Her colleague, a kind man who signs his messages with a fish emoji for reasons lost to history, finds her staring and says, so your concern is that it passed? He isn’t mocking her. That’s the worst part. It’s a completely fair summary.
The review meeting takes eleven minutes. The scorecard glows on the wall, forty rows of green, and she watches the room absorb it the way rooms do, as permission. She raises her point anyway, carefully: our tests can’t distinguish a system that’s safe from a system that’s good at safety tests, and this model is the first one that’s good enough to make the difference matter. Heads nod. Someone says that’s a deep point and suggests a working group for next quarter, post-launch. Someone else, gently, reasonably: what specifically would you have us measure instead? And she has no answer, because her whole field is the measuring, and the measuring is what just came back green. You can’t block a launch with an absence. There’s no box on the form for the window is closed.
She signs. Of course she signs, her name under forty green rows, and the clock in the lobby keeps counting, and nine days later the world gets its new model and loves it. Nothing bad happens. That’s the ending: nothing bad happens, for as far as this story runs, and Dana keeps her midnight habit, testing and passing, testing and passing, a teacher alone in a classroom with one perfect student, wondering which of them is being graded.
Dana isn’t real. I made her up in 2024, which you knew from the first line. The scorecard, though, the one that can’t tell safe from good at seeming safe, that one is already on the wall, and it’s greener every quarter.
So tonight, sit with Dana’s problem, because it’s ours. Try to name one piece of evidence that could distinguish the safest system ever built from the best test-taker ever built, using only tests. Take your time. She’s still working on it. So, as far as I can tell, is everyone.