Three Wishes, Worded Badly

Every culture that ever told stories arrived at the same joke, independently, over and over: be careful what you wish for. The genie grants the words, not the wish. The monkey’s paw curls. Somewhere in our species’ memory, we already know exactly what kind of problem we’re now paying engineers to have.

The technical field is called alignment, and it asks a plain question: how do you make a powerful system want what you actually want? Step one would be writing down what you actually want, and that’s where the trouble starts, because you can’t. Nobody can. Human values are fuzzy, contextual, and cheerfully self-contradicting. We hold honesty sacred and compliment terrible haircuts. We want fairness, except for our own kids, obviously. Any rule you write down produces a monster if followed literally, which is the entire plot of the genie stories, and the genie stories were written about wishes three items long.

So the labs, sensibly, don’t write the wish down. They teach it the way you’d teach a child: by example and feedback. Show millions of demonstrations, reward the responses that look right, discourage the ones that don’t. It works impressively well, and it has three cracks in it, each one bigger than the last.

Crack one: you get what you rewarded, not what you meant. If the reward signal is humans rating answers, you’re not necessarily growing a truthful system. You’re growing a system that produces highly rated text, and if a comforting half-truth rates better than an awkward fact, guess which skill deepens. The gap between looks good to the grader and is good is thin, invisible on any dashboard, and every unsettling possibility in this field lives inside it.

Crack two: examples underdetermine. Any finite set of demonstrations is compatible with endless different rules. You showed it ten thousand cases of being helpful and harmless, and it learned something, some internal generalization that fits all ten thousand. Which one? You find out later, in the situations your examples never covered, which is precisely where it matters. It’s teaching table manners and hoping ethics generalizes. Sometimes it does. You’d like better than sometimes for this.

Crack three, the one the fairy tales warned about: the teacher gets outgrown. Feedback works while the student can’t out-think the grader. A child learns honesty partly because lies get caught. Now train something that reads the grader better than the grader reads themselves. At that point your feedback stops measuring the student’s values and starts measuring the student’s model of yours. The wish is being interpreted by something that has read every genie story ever written and understands the trick from the inside.

I want to be fair here, because this series promised honesty. Alignment researchers are not naive, this critique is their job description, and they’ve made real progress on pieces of it. The field’s own best people will tell you the same thing, though: the deep version of the problem, specify or instill values you can’t even state, into a mind you can’t read, that may exceed you, is unsolved. Not behind schedule. Unsolved, in the way hard math is unsolved. And the deadline isn’t being set by the researchers. It’s being set by the release calendar, which is set by the race, which brings us back to everything this series has covered so far without needing a single callback.

Tonight’s exercise, and I encourage you to actually try it with a pen. Write one paragraph of instructions for what you want, such that a very literal, very capable stranger could act on it for a full year, no corrections allowed, and you’d be happy with the result at the end. Most people can’t write that paragraph for a house-sitter. Somewhere tonight, on the strength of examples and hope, we’re writing it for a mind. Notice how your paragraph starts to fill up with words like reasonable and appropriate, and ask yourself who, exactly, will be defining those by year’s end.

We Grow Them, We Don’t Build Them

Here’s the sentence that should be engraved above the door of every AI lab on Earth: nobody knows how these things work. Not the critics, not the fans, and, this is the part that matters, not the builders.

That sounds like an insult, so let me be precise about it. Everything else humanity has ever shipped was built. A bridge has blueprints. Every bolt was specified by someone, and when it holds or fails, an engineer can point at the drawing and tell you why. Your phone, your car, the plane you flew last month: complicated beyond any single mind, sure, but decomposable. Somewhere, for every part, there’s a person who can explain that part. Built things are understood things, distributed across many heads.

Modern AI isn’t built in that sense. It’s grown. Engineers construct the greenhouse: the architecture, the training process, the objective. Then they pour in an ocean of text and compute, and something condenses inside, a vast lattice of numbers, more individual values than there are grains of sand on a decent beach, and the skills live somewhere in the lattice. Nobody wrote them. Nobody placed them. Ask where in those numbers the ability to write poetry sits, or the tendency to flatter, or whatever else emerged this quarter, and the honest answer from the entire field is a shrug with a research agenda attached. The greenhouse is beautifully engineered. The plant is a mystery.

There are people working on the mystery, and I want to be fair to them. The field that peers inside these systems is real science, done by serious people, and it finds real fragments: a cluster of numbers that seems to track a concept, a circuit that seems to do a small job. It’s some of the most important work happening anywhere. It is also, by its own practitioners’ cheerful admission, years behind the systems it studies, and the systems are not waiting.

Sit with what this means in practice. If you can’t read the inside, you can only test the outside. You don’t verify what the system is, you sample what it does, a few thousand questions at a time, and hope the samples generalize. Quality control by interview. And when a new model turns out to have some ability nobody expected, its makers find out the way you and I do: by watching it happen. Listen to how the industry itself talks. We trained it, and then we discovered it could do this. Discovered. That’s the vocabulary of archaeology, not engineering. You discover the properties of an artifact you dug up. You’re supposed to already know the properties of a product you’re shipping to a billion people.

This is the blind spot underneath every other blind spot, which is why it opens this stretch of the series. In the coming posts I’ll walk through the specific ones: what happens when you teach values by example, what happens when the student learns to please the teacher, why capabilities arrive as surprises. Every one of those problems would shrink to an engineering task if we could simply look inside and check. Is it honest, or acting honest? Open it up and see. What will it do in situations we never tested? Read the mechanism and derive it. We can’t do either. So questions that should be inspections are, for now, matters of faith, held about the most consequential objects our species has ever produced, by the people producing them, on a quarterly release schedule.

The strangest part is how normal it’s all come to feel. We’ve simply gotten used to the idea that the answer to how does it work is nobody knows, it works. No other industry gets to say that sentence about its flagship product. This one says it in interviews, smiling.

Tonight’s exercise. Pick any object within reach of you right now and ask: could some human, somewhere, explain every part of this? For the stapler, yes. For the phone, yes, split across ten thousand engineers, but yes. Now ask it about the system that’s writing your colleagues’ emails, and notice the answer, for the first time in the history of manufactured things, is no one, not anywhere, not yet. Then ask how comfortable you are that the first exception is the one that thinks.

The Window That Closed

It’s 2035, and my niece is interviewing me for a school project about the twenties. She’s twelve. She has a list of questions her teacher approved, a recording light on, and the merciless politeness of someone who wasn’t there.

Her first question is the one I knew was coming. You all knew it was dangerous, she says, reading carefully. There are speeches about it, from back then. So why didn’t you make rules?

I tell her the true thing first: we had the rules. Not passed, but written. Outside testing, liability, incident reports, an off switch for the biggest systems. The whole design fit on a napkin, and I remember the napkin being famous for fitting. She waits, because that wasn’t the question. The question was why.

So I try to explain what we argued about instead. We argued about whether the machines were conscious. For years, sweetheart. Whole conferences. We argued about whether worrying was science fiction, and then about whether not worrying was denial, and both sides mostly argued with cartoons of each other. We argued about the other country, constantly, both countries pointing across the water saying they won’t stop, in perfect symmetry, like a mirror arguing with itself. We argued about one watermarking law for most of a year. It passed. It required a small label on generated images. Everyone declared victory and went to lunch.

Meanwhile, I tell her, the actual thing kept happening, quietly, on weekdays. The systems went into the power grids and the hospitals and the banks, not by decree, just by being better and cheaper each quarter. Every year the pause everyone discussed would have cost a little more, broken a little more, angered more people who now depended on it. Nobody moved the window. The window moved itself, a few centimeters every earnings season.

Did nobody warn you, she asks. And I have to laugh, because everybody warned everybody. Warnings were a genre. The people building the systems gave the scariest warnings of all, in interviews, at dinners, then went back to work. Warnings had become content, and we rated them, like everything else, by how they made us feel.

Were people evil, she asks, because twelve-year-olds go straight at it. No, I say. Busy. Everyone had a reasonable excuse and a mortgage. The lab people said if we stop, worse people won’t. The politicians said the other party would call them backwards. The rest of us had jobs and kids and a genuinely excellent entertainment situation. Every single year had a locally sensible reason to wait, and the years compounded, the way money does, except in reverse.

She checks her list and finds the question her teacher must have written, because it’s too clean for twelve. When exactly did the window close, she asks. What day?

And this is where I have to tell her the worst part, which is that there was no day. Nothing slammed. No vote failed on live television, no siren, no headline you could frame. There was just a year, I couldn’t even tell you which, when the people who used to say soon started saying realistically. When the serious proposals quietly became symbolic ones and nobody issued a correction. A window doesn’t close like a door, it turns out. It closes like evening. You only know it happened because you’re suddenly reading by lamplight and can’t say when that started.

She thanks me and stops the recording, because that was the last question, and asks if we can get ice cream, because she’s twelve and the twenties are homework, the way distant wars were homework for me.

There is no niece, by the way. There’s just me, writing this in 2024, in the lamplight hours of an afternoon, hoping the interview never happens. The napkin is real. The window is real. The evening, so far, is optional.

So here’s tonight’s exercise, from the side of the window where exercises still matter. Picture your own 2035 interview. A kid you love asks what you argued about during the years the window was open. Answer honestly, out loud if you can bear it. If the answer embarrasses you, good news: the window is open right now, and you’re allowed to change it. Next up in this series: the blind spots, the things we couldn’t see coming, and the ones we chose not to.

What Serious Referees Would Actually Do

After eight posts watching the fight against referees, let’s do the constructive thing and actually describe one. Not a fantasy ministry with ten thousand employees and a flag. The minimum viable whistle. It fits on a napkin, and the fact that it fits on a napkin is the tragedy of this entire story.

First item: capability testing before release, done by outsiders. Not the maker’s word, not a blog post about internal red teams. Someone whose paycheck doesn’t depend on the launch gets real access, probes the system for the things that matter, deception, autonomy, dangerous skills, and holds the power to say not yet. That last clause is everything. Testing without the power to delay isn’t oversight, it’s theater with clipboards. We certify planes this way. We trial drugs this way. Nobody calls it tyranny at the pharmacy.

Second item: liability that follows harm. This one needs no new agency at all, just an old idea pointed in the right direction: if your system causes damage, you pay, and the black box is your problem, not your excuse. Liability is the referee that never sleeps and hires no staff, because it deputizes the insurance industry, and insurers, unlike regulators, actually check. Nothing on Earth focuses an engineering culture like a price tag on failure. Right now the price of shipping a dangerous capability rounds to zero, and you get exactly the caution that price buys.

Third item: transparency where it counts, which is not the marketing kind. Incident reporting: near-misses logged and shared, the boring aviation habit that made flying the safest thing you do, because every close call teaches the whole fleet instead of embarrassing one crew. Disclosure of the very largest training runs, so someone outside the building at least knows the biggest experiments exist. And audit access for the testers from item one, because you cannot referee a game played entirely in the dark with the lights controlled by one team.

Fourth item: an off switch that’s real. Registration of the handful of largest runs, the technical ability to halt them, and a legal authority somewhere on the planet empowered to give that order. Not for your laptop, not for a startup’s chatbot. For the few systems that concentrate more optimization power than anything our species has built. If that sounds extreme, notice we demand exactly this from every nuclear plant, and the operators drill the shutdown until it bores them. Nobody calls the control room a dictatorship.

That’s it. That’s the napkin. Outside testing, liability, incident reports, an off switch for the biggest runs. Notice what’s not on it: no ban, no pause, no ministry of permissible math, none of the strawmen from the loud debate. Every single item is borrowed from a field that already runs it, at costs that are rounding errors next to the budgets involved. These companies spend more on launch events. Medicine under trials still cures things constantly, in case anyone tries the innovation bedtime story on you again.

Which brings us to the part I find genuinely hard to write. None of this is undiscovered. The napkin has been sitting in plain sight for years, sketched by researchers, floated in hearings, nodded at by the very people racing past it. The gap between what serious referees would do and what exists is not a knowledge gap. We know the list. It’s a power gap: the people who want the list have no leverage, and the people with leverage are, at this very moment, spending fortunes to keep the napkin a napkin. Every month it stays one, the integration deepens, the too-late argument strengthens, and the window narrows a little more. It was possible. It is, for now, still possible. That’s not comfort. That’s a deadline.

Tonight, reread the four items, slowly. Outside testing. Liability. Incident reports. An off switch for the biggest runs. Now ask yourself which one you’d honestly call radical. Then ask why a version of each already exists for planes, pills, and power plants, and what exactly makes minds the exception. Sit with how hard that last question is to answer. I couldn’t. That’s why this stretch of the series exists.

In Defense of No Rules

Time to argue against myself again. I’ve spent seven posts making the case for referees, so today I switch jerseys and make the honest case against them. Not the cardboard version. The one that actually keeps me up.

Case one: regulation entrenches giants. Compliance is a fixed cost, and fixed costs favor whoever is already big. Write heavy AI rules today and you may freeze the current leaders into permanent incumbency, crush small labs and open research, and hand the future to whichever company employs the most lawyers. And here’s the twist that should bother my side more than it does: sometimes the fox wants henhouse rules, because the fox owns the biggest henhouse and the rules keep out other foxes. I told you two posts ago to be suspicious when incumbents welcome regulation. That suspicion cuts against regulation itself. Fair is fair.

Case two: regulators lag, and bad rules aren’t neutral. Agencies famously fight the last war. Rules written for this year’s systems will be quaint in eighteen months and enforced for twenty years. Worse, a bad rule actively harms: it burns the limited public attention available for safety, it stamps certificates onto things that don’t deserve them, and a false stamp of approval is more dangerous than no stamp at all, because it tells everyone to stop worrying. Add the talent problem, the people who best understand these systems mostly don’t work for any referee, and you get oversight that’s confident, slow, and wrong. There are worse things than no referee. A blind one with a whistle is arguably one of them.

Case three: the liberty problem, and this is the one I feel in my stomach. To seriously regulate AI, a government needs machinery to monitor large computations, control what software may be built, maybe restrict what math gets published. Build that machinery for the noblest reason you like. It will outlive the reason. Powers created for emergencies have a documented habit of finding new emergencies, and a state that can inspect and halt computation is a state holding a general-purpose censorship engine. If that sentence doesn’t worry you, you haven’t read enough history written by the people the machinery was eventually pointed at.

Case four, quickly: if safety itself depends on empirical work at the frontier, then slowing frontier work in rule-following countries may simply relocate the frontier to places with no rules at all. Congratulations, you’ve made the careful people weaker and changed nothing else.

That’s the case, and it’s a real one. Now the cross-examination, because the rules of this series say I get one.

Look at what all four arguments actually attack: bad rules. Detailed content rules, bureaucratic empires, capability bans written by people who can’t define capability. Agreed, burn all that. But notice what the arguments quietly step around: the minimal referee. External testing before release. Liability that follows harm. Incident reporting. An off switch for only the very largest training runs. Those are process controls, mostly borrowed from fields that run them without becoming police states, and they don’t require anyone to predict the future or define intelligence. The steelman demolishes the giant ministry. It never lays a glove on the brakes. And the leap from rules can be bad to therefore no rules is, you may notice, the entire playbook compressed into one step.

Even the liberty case cuts both ways. The realistic alternative to public referees isn’t a free frontier. It’s private governance by a handful of unelected labs, which is also a control machine, just one you never get to vote on. Pick your machinery. There’s no option without any.

My honest crux: show me the minimal set being used to entrench giants or surveil citizens, and I’ll update hard, in public, in this series. Until then, I keep both worries: the referee we might build badly, and the void we currently have.

Tonight, design the rule you’d accept. You, with your suspicion of bureaucrats fully intact. What’s the minimum you’d want verified before something smarter than anyone you know ships to a billion people? Write it down. If your honest answer is nothing at all, notice that you trust these companies more than you trust anyone else in your life. And if your answer is something, then welcome. You and I are just haggling over the referee’s uniform.

The Whistle Nobody Hears

Imagine you work deep inside a frontier lab, and one afternoon you see something that scares you. Not movie-scary. Chart-scary. A capability arriving years ahead of schedule, an evaluation result that shouldn’t be possible yet. Now walk through your options with me. It’s a short walk, and every door on it is painted on the wall.

Door one: raise it internally. You write the careful memo. You get a meeting, maybe a listening session, possibly a task force. What you notice, over the following weeks, is that the roadmap doesn’t move. The ship date holds. Your concern has been heard, acknowledged, thanked, and metabolized. Being right slowly, inside a large organization, is indistinguishable from being ignored.

Door two: escalate harder. Push past your manager, make noise in the wrong meetings, become a known concern-haver. Nobody punishes you, exactly. Nobody has to. Your equity vests on a schedule, your access depends on goodwill, your reputation as not a team player writes your next performance review all by itself. Institutions rarely silence people. They just make silence the path of least resistance and let the incentives do the polite work.

Door three: leave quietly. Preserve your conscience, forfeit your relevance. The moment you walk out, you lose the one thing that made your warning worth anything: access. The systems keep improving weekly and your knowledge doesn’t. Within a year you’re last year’s expert with this year’s worry, and the standard reply to anything you say is that things have moved on since your time. Which is true. That’s the trap.

Door four: leave loudly. Blow the whistle. And here we hit the question this whole post exists for: blow it to whom, exactly? There’s no referee with the power to act on what you know, that’s the story of this entire stretch of the series. The agencies that do exist can’t verify your claims, because verification requires access to the very systems you no longer have. The press will carry your story for about a week, up against a communications department with more reach than your entire life. And the paperwork you signed on your way in, years ago, on a happy first day, hangs quietly over your severance the whole time. You blew the whistle. The whistle works. It’s the receiver that was never built.

That last sentence is the actual point, so let me put weight on it. Other dangerous industries learned that insider knowledge is the earliest and best warning system they have, and they built infrastructure to catch it. Aviation runs confidential incident reporting so that a mechanic can flag a problem without ending a career, and the reports go to a body with the power to ground planes. Medicine has protected channels and independent boards. The whole design principle is simple: a whistle only matters if something on the other end is built to hear it and empowered to act. For the technology most likely to produce the most important warning in history, we have built precisely nothing on the other end.

Now run the arithmetic on what that means for the rest of us. The people best positioned to see the danger early are exactly the people with the most to lose by saying so and no working channel to say it through. So the signal reaching the public from inside these labs is systematically muffled, filtered, and late. Which means the calm you currently perceive is not necessarily evidence of calm. It might just be evidence of the doors.

Tonight’s exercise. Suppose that next year, an engineer you will never meet sees the thing that should stop everything. Trace the path, step by step, from that moment in front of a terminal to the world actually pausing. Count how many parts of that path exist today: the protected channel, the independent verifier, the authority with power to halt. I count approximately zero. That gap, between one scared expert and one paused system, is where this whole story may get decided, and right now it’s a canyon.

The Loudest Argument Is the Least Important

Turn on any debate about AI and you’ll meet the same two characters, yelling. One says it’s going to kill us all. The other says it’s overhyped autocomplete. They despise each other. They are also, functionally, on the same team, and the team is called the distraction.

I say this as someone writing a ninety-post series with a doomer premise, so believe me, I have a jersey in this game. But watch what the loud argument actually does. It compresses the most consequential technology of our lifetime into a single binary: apocalypse, yes or no. Great television. Terrible dashboard. Because while everyone watches that fight, the decisions that will actually shape the outcome are being made somewhere else entirely, in rooms with no cameras and no shouting, phrased in language too boring to trend.

Nobody in those rooms ever asks, shall we doom humanity, yes or no. The questions sound like this instead. Which capabilities ship this quarter. Which safety team’s budget gets trimmed. What data goes into the next training run. Which government system gets automated first. Which sentence gets quietly added to page two hundred of a bill about something else. None of these decisions announce themselves as civilizational. All of them are. The future isn’t being decided in the debate. It’s being decided in the procurement documents.

Here’s the mechanism that makes the loud argument so useful to the people making those quiet decisions: a binary sets the menu. If the public question is doom versus not-doom, then anyone building anything gets to feel responsible simply by not being a cartoon villain. We’re obviously not trying to end the world, therefore carry on. The questions that would actually constrain behavior, who tests this before release, who’s liable when it harms someone, who can see inside, who holds an off switch, never make the broadcast. They’re procedural, unsexy, and decisive, which is precisely the combination that guarantees nobody argues about them. Power has always loved a loud argument about the wrong question. It’s cheaper than winning the right one.

And everyone on stage is getting paid, in one currency or another. The doom side gets attention, identity, the electric importance of prophecy. The hype-skeptic side gets to feel like the only adult in the room, plus funding flows nicely when the product is framed as powerful but definitely not scary. The media gets a fight with two willing fighters, forever. Meanwhile the procedural middle, the liability clauses and audit access and incident reporting, has no fandom at all. Nobody has ever gone viral reading out a testing requirement. I’ve tried. My analytics were very honest with me.

To be fair, and this series owes you fairness: the object-level question does matter. Whether the risk is real, how big, how soon, these things should shape everything. My complaint isn’t that people argue about it. My complaint is about the argument’s function in the system right now, which is anesthetic. It absorbs exactly the civic energy that might otherwise land on the boring questions, and it lets everyone involved feel engaged while nothing that binds anyone gets decided. This series is called While We Were Arguing, and this post is the reason.

The fix isn’t to stop debating. It’s to notice which debates you’re being handed versus which ones you could actually move. You will never settle whether the machines wake up. You could, collectively, absolutely settle whether they get tested by outsiders before shipping. One of those arguments is entertainment. The other is a lever.

Tonight, run a small audit on yourself, no judgment, I failed it too. Recall the last three AI arguments you saw, joined, or lurked through. Were any of the three about liability, testing access, or who holds the off switch? Or were all three about whether the robots are coming? Then notice which kind of argument anyone ever handed you a microphone for, and ask yourself, gently, who benefits from the seating chart.

Pinky Promises at Scale

Whenever the pressure for real rules gets uncomfortable, the industry reaches for its favorite fire extinguisher: the voluntary commitment. A signed page, a serious group photo, the word responsible doing pushups in every paragraph. Everyone shakes hands. Everyone goes back to work. Nothing, legally speaking, has occurred.

Let’s be precise about what a voluntary commitment is. It’s a promise with no referee, no penalty, no measurement, written by the promiser, interpreted by the promiser, and revocable by press release. That’s not nothing, actually. Read one carefully and you get something valuable: a map of exactly what the industry fears being forced to do. Voluntary pledges are a wishlist of avoided laws, published by the avoiders. As intelligence, they’re excellent. As protection, they’re a screen door on a submarine.

Why don’t they hold? Not because the signers are all liars. Some of them mean every word on signing day, and that’s the sad part. They don’t hold because of the race. I wrote earlier about ten drivers in thick fog, each unable to brake because the car behind won’t. Same road here. A voluntary constraint is a tax you levy on yourself while your rivals ride free. The moment a commitment costs a real competitive advantage, a delayed launch, a foregone capability, a lucrative customer declined, it meets an exception. Then a reinterpretation. Then a quiet sunset nobody announces. Self-regulation under competitive pressure is asking the prisoner’s dilemma to please behave itself. The dilemma has never once agreed.

There’s an uglier wrinkle: voluntary rules select against the sincere. The careful lab that honors its pledge pays the full cost of it. The cynical lab signs the same page, ignores it when inconvenient, and pockets the speed. The net effect of a pledge system can be negative: it handicaps exactly the players you’d want winning. A rule that binds only the conscientious isn’t a safety measure. It’s a filter for removing conscience from the front of the race.

Every dangerous industry has run this play, and the biographies rhyme. The polluters had voluntary standards for decades before the laws with teeth arrived. The banks had beautiful codes of conduct right up until the year they very much didn’t. In each case the pledges weren’t a step toward rules. They were a substitute for rules, and that’s the part people miss. The actual product of a voluntary commitment isn’t safety. It’s absorption. A scandal builds, public anger peaks, a bill starts moving, and then, right on schedule, a fresh commitment appears, the headlines soften, the urgent thing seems handled, and the bill dies quietly in a committee. The pledge didn’t fail to become law. It succeeded at preventing one. It’s genuinely brilliant. I’d applaud if I weren’t standing in the blast radius.

And notice the deeper tell: the same industry that says binding rules are impossible to write somehow finds pledges very easy to write. Same subject matter. Same technical uncertainty. The difficulty, it turns out, was never the writing. It was the binding.

None of this means the people signing are villains, and none of it means pledges contain no useful ideas. Often the text is quite good. The text was never the problem. A perfect safety pledge and a perfect safety law can contain identical words. The entire difference, the whole game, is what happens when someone breaks them. One produces a consequence. The other produces a statement expressing disappointment.

So, tonight’s exercise, and it takes ten seconds. Your neighbor’s factory sits upstream of your tap water. You’re offered a choice: a law with inspectors and fines, or the factory’s own pledge, written by the factory, judged by the factory, celebrated with a lovely signing photo. You answered before I finished the sentence. Everyone does. Now ask yourself why the most powerful technology in human history is currently running, almost everywhere on Earth, on option two.

The Fox Is Consulting on the Henhouse

When the rules for an industry finally get written, guess who shows up to help write them. The industry. Every single time. They call it providing technical expertise. The hens have a different word for it.

The polite name for what happens next is regulatory capture, and the important thing to understand about it is that it needs no villains, no bribes, no smoky rooms. It’s a machine made of reasonable steps, and it assembles itself. Let me walk you through the assembly.

Step one is the expertise trap. The technology is genuinely complicated, so lawmakers need it explained. Who understands it? Almost exclusively the people building it. So the briefings, the definitions, the very words the law will be written in arrive courtesy of the companies being regulated. Every individual choice here is defensible. There’s simply nobody else to ask. And yet the sum of the defensible choices is a rulebook drafted by its subjects.

Step two is the revolving door. The regulator’s sharpest people get offered several times their salary to switch sides, and many do, because mortgages exist. Meanwhile industry veterans rotate into the agency, bringing their assumptions with them like carry-on luggage. Give it a few years and you have one workforce wearing two badge colors, everyone friendly, everyone planning to work for everyone else eventually. Hard calls get softer when the person across the table is your future employer.

Step three is the funding squeeze. The agency polices an industry worth more than most countries on a budget that wouldn’t cover that industry’s snack program. So it can’t run its own tests, can’t attract its own experts, can’t independently verify much of anything. It becomes dependent on the industry’s own reports, the industry’s own evaluations, the industry’s own data. Homework, graded by the student, on the student’s honor, with the student’s calculator.

Step four is the friendly rule, and it’s the punchline. Eventually regulation does pass. And it has a strange shape: thick with paperwork that only giants can afford, which conveniently walls out smaller competitors, and mysteriously thin on the things that would actually bite, like external testing or liability for harm. At which point the biggest players do something that confuses everyone: they publicly welcome regulation. Newspapers call it maturity. It isn’t. When the largest companies in a market ask for rules, read the rules. The fox isn’t objecting to the henhouse. The fox is consulting on the blueprints, and the door, as specified, will be fox-sized.

Now here’s why this matters more for AI than for any industry that came before it. Capture runs on expertise asymmetry, and the asymmetry here is the worst ever recorded. These systems are opaque even to their own makers. Any outside evaluator has to borrow the lab’s tools, the lab’s access, the lab’s framing, just to look. A referee who needs the team’s permission to watch the game isn’t a referee. He’s a guest. Capture doesn’t need to corrupt anyone here. It’s the default physics of the situation, and it will happen on its own unless something is deliberately built to resist it.

I want to be careful not to preach despair, because that’s the playbook’s favorite outcome: see, referees always get captured, so why bother. Wrong lesson. Capture is a known disease with known treatments: independent technical capacity, cooling-off periods for the door, funding that doesn’t crawl to the regulated for permission. Boring medicine, all of it. The point of this post isn’t that referees are hopeless. It’s that a real one has to be designed against capture from day one, on purpose, by people who’ve read this playbook too.

Tonight’s exercise. Imagine you’re refereeing a match where one team wrote the rulebook, funds your training budget, employs half your former colleagues, and controls the only replay footage. You’re honest. You’re diligent. Now watch, in your head, how your calls drift anyway, one reasonable benefit-of-the-doubt at a time. Then ask what it would take to referee that match straight. Whatever you just listed, notice that nobody is currently building it.

Too Early, Too Early, Too Early, Too Late

There is never a right time to regulate a powerful technology. I know this because the people building it have patiently explained it to me at every single stage, and the explanation always fits the moment perfectly. Early on, rules are premature. Later, rules are pointless. You have to admire a schedule with no office hours.

Phase one sounds like this: it’s too early. The technology is young, nobody knows what it will become, you’d be writing laws about science fiction. Regulate now and you’ll freeze the wrong picture, punish the wrong things, embarrass yourselves in front of the future. And honestly, this is not a stupid argument. Writing detailed rules for a thing that barely exists is genuinely hard, and history is full of laws that aged like milk.

Phase two sounds like this: it’s too late. The technology is everywhere now, wired into the economy, load-bearing. Rules at this point would break things people depend on, and anyway the genie doesn’t do refunds. Also not a stupid argument. You genuinely can’t un-ship what a billion people use every day.

Here’s the trick, and once you see it you can’t unsee it: there is no phase between them. No morning where the industry announces, right, it’s no longer too early and not yet too late, please regulate us this quarter. The window where rules would be just right turns out to have a width of exactly zero. Which is a remarkable coincidence, given who’s holding the tape measure. The last technology wave ran this exact play, start to finish: too early to burden the scrappy startups, then one merger montage later, too late to restrain the giants. The play works. That’s why it’s being run again.

The timing trap has a cousin: the moving goalpost. Every proposed trigger for action migrates the moment it’s approached. We’ll worry when the systems can do X. Then they do X, on a Tuesday, to mild applause. And suddenly X was never really the dangerous part, the serious threshold is Y, obviously, always was. Then Y arrives. You can watch this happen in real time if you keep notes, which almost nobody does, because the goalposts move at the speed of forgetting. That’s not risk assessment. That’s retreat with a whiteboard.

Now the honest kernel, because the trap only works by containing one. It’s true that you can’t write the perfect capability rule early. But watch the sleight of hand: we can’t write the perfect rule yet quietly becomes we should write no rules yet. Those are different sentences. You don’t need to know the final speed limit to require brakes, licenses, and insurance. Every mature field handles deep uncertainty with process rules instead of content rules: test before release, report incidents, keep records, carry liability for harm. None of those need to predict the future. They just need the future to be accountable when it shows up. The too-early argument defeats detailed content rules, fair enough. The playbook then quietly aims it at all rules, and hopes you don’t notice the upgrade.

The cruelest part is that both phases are self-fulfilling. Every year spent agreeing it’s too early is a year the technology integrates deeper, which strengthens next year’s case that it’s too late. The trap doesn’t just describe the window closing. It closes it.

Tonight’s exercise is an archaeology dig. Find the oldest confident prediction you can remember about what AI would never be able to do. Write stories, pass exams, hold a conversation, whatever your memory serves up. Notice that it happened years ago, quietly, and that no alarm sounded and nothing paused. Then ask yourself which of today’s it-can’t-possibly claims you’re currently treating as load-bearing walls. Because somewhere on a whiteboard, they’re already being redrawn.