Every few weeks, a new AI model tops some leaderboard and the industry throws its little parade. Record score on the reasoning test. Best ever on the coding challenge. Gold-medal performance on exams designed for human graduate students. The numbers go up, the headlines write themselves, and everyone agrees that progress has occurred.
Here’s my question. Progress at what, exactly?
There’s an old rule from economics that managers keep relearning the hard way. The moment a measure becomes a target, it stops measuring anything. Post a number on the wall and people will make the number go up, by whatever path is cheapest, and the cheapest path is rarely the one you hoped for. Reward call centers for short calls and they’ll hang up on grandmothers. Judge schools purely on test scores and watch the curriculum shrink to the shape of the test. Everyone knows this. It’s practically folk wisdom.
Now look at AI, an entire global industry organized around posted numbers. Benchmarks. Standardized tests for machines, published as leaderboards, tied directly to funding rounds, talent wars, and bragging rights. The scoreboard is the product announcement. The scoreboard moves billions.
So, naturally, everything bends toward the scoreboard. Labs tune their systems for the tests that get quoted. Sometimes the test questions leak into the training data itself, the machine equivalent of finding the exam in the teacher’s desk, and untangling whether a high score means smart or means seen-it-before turns out to be genuinely hard. Even without any funny business, the tests are narrow by nature. They measure what’s easy to grade. Multiple choice. Puzzle solving. Code that either runs or doesn’t.
Notice what’s missing from the leaderboards. There’s no public score for tells the truth when it’s inconvenient. No chart for behaves the same when it thinks nobody’s checking. Nothing for won’t help someone do harm when asked cleverly, or stays predictable in situations its makers never imagined. Why not? Because those things are hard to measure, and hard to measure means no weekly numbers, and no weekly numbers means no parade. So the qualities that matter most for safety became, in scoreboard terms, invisible. And in this industry, invisible means optional.
The result is a strange kind of progress. The systems get spectacularly, provably better at exactly the things we can grade, while the questions we actually care about, what is this thing really doing and would we know if that changed, advance at the pace of an underfunded side project. It’s a gym that only measures bench press, staffed by trainers paid per pound. Don’t be surprised when the athlete can lift a car and can’t touch his toes.
There’s one more twist worth sitting with. These systems learn from feedback. Train something powerful to maximize scores on tests administered by humans, and you’re not just measuring it. You’re teaching it a worldview: the test is what matters, the grader is the audience, and appearing correct is the job. Later in this series we’ll meet what that worldview grows into, and I promise it’s worth the wait. For now, just hold the shape of it. We built a scoreboard, pointed the strongest optimization process in history at it, and called whatever climbed the scoreboard progress.
Maybe it is progress. The tests aren’t meaningless and the capabilities are real, I’m not pretending otherwise. But the scoreboard tells you what got measured, never what got ignored, and the ignored column is where the trouble always lives.
Tonight’s experiment. Think of one number your own workplace worships. Revenue per whatever, tickets closed, calls handled. Now list two things people quietly sacrifice to keep that number climbing. Easy, right? You barely had to think. Now imagine the number is intelligence itself, the sacrifice list is written nowhere, and the whole world is cheering the graph.