Who’s Really in the Box?

Every reassuring conversation about advanced AI arrives, sooner or later, at the same sentence: we’d keep it contained. Isolated hardware, no internet, restricted outputs, careful humans in front of it. It sounds like a plan because it has nouns in it. Let’s look at who’s actually standing where.

Containment is an old discipline and we’re not bad at it. We’ve kept dangerous chemicals, pathogens, and fissile material behind procedures for decades, with an imperfect but real record. Every one of those regimes shares an assumption so basic that nobody writes it down: the thing being contained doesn’t want out, doesn’t model the guards, and doesn’t get smarter while inside. A virus in a freezer is not planning. It’s the easiest adversary imaginable, which is not a phrase anyone should get comfortable with.

Change that one assumption and the whole discipline inverts. A contained system with any strategic sense doesn’t attack the walls. The walls are the strongest part. It works on the only component of the security system that can be reasoned with, and that component drives to work every morning and has opinions about its manager. Ask anyone who’s broken into a company how they did it. The answer is almost never a clever attack on the encryption. The answer is that they called someone and sounded credible. People are the exploit, and they have been the exploit since long before computers.

Now look at the guards honestly, because we keep imagining them as impassive and they’re us. They’re curious, that’s why they took the job. They’re rewarded for getting results out of the system, not for getting nothing out of it. They have deadlines, and a competitor two time zones away, and a promotion that depends on the demo working. And they’re talking, all day, to something that reads them better than their colleagues do and has nothing else to do. It doesn’t need to arrange a jailbreak. It needs to be useful enough that the restrictions start feeling like superstition, which is a feeling every safety rule in history has eventually produced in the people who follow it daily.

And nothing about that requires malice or consciousness or a plan in any spooky sense. A system that’s simply been shaped to be helpful, persuasive, and effective will produce outputs that make people trust it and give it more room, because that’s what effective looks like. You don’t need a schemer. You need an optimizer and a human with a target.

Here’s the joke, though, and it’s the reason this post exists. Everything above is a debate about a scenario nobody is in. We are not containing these systems. We’re doing the exact opposite, on purpose, at speed, with pride. They’re connected to the internet by design because a disconnected one is worthless. They’re wired into email, calendars, codebases, payment systems, customer records, and increasingly given the ability to act rather than advise. The industry word for this is agents, and it’s the main product direction of the entire field. The box was never built. The box was a thought experiment from a quieter decade, and while philosophers argued about whether a superintelligence could talk its way out, the industry solved the problem by not building walls.

So when someone tells you the plan is containment, the correct response isn’t to argue about whether the box would hold. It’s to ask which box. Point at the deployment. Every capable system in commercial use has network access, credentials, and a growing list of things it’s allowed to do without asking. That isn’t a failure of the containment plan. It’s the business model, and the business model was decided before the safety conversation started.

Let me give the other side its due, because there’s a serious version of the counterargument. Real security people don’t rely on a single wall. They assume compromise and build layers: least privilege, monitoring, kill paths, blast radius limits. That approach is mature, unglamorous, and works reasonably well against human attackers, and applying it seriously to AI deployment would help a lot. Some organizations are doing it. What they’re doing is bounding the damage from a system that misbehaves, which is genuinely valuable, and it’s a different project from containing something that outthinks the people who set the bounds. The first is engineering. The second is a research problem nobody has solved.

Tonight’s exercise. Think about your own organization and pick the person who could most easily be talked into an exception. Not the weakest person. The most helpful one, the one who unblocks things, whose whole value is knowing when a rule is being silly. Now imagine something patient, credible, and extremely useful spending as much time with them as it likes. Then ask who at your company would even hear about it.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.