The Thermostat Doesn’t Hate the Cold

The most common reassurance about all of this is that these systems have no feelings, no desires, no self. It gets offered as though it settles the matter. It’s true, as far as anyone knows, and it’s an argument for the other side, and today I want to explain why as plainly as I can.

Start with a thermostat. It has a goal in the only sense that matters to the room: it acts to bring about a state of the world and it acts against things that move the world away from that state. It doesn’t hate the cold. It has no opinion about winter, no ambition, nothing it’s like to be. And if you open the window it will fight you, all night, without emotion and without tiring, and if you want the window open you’ll have to deal with the thermostat as a fact rather than reason with it as a person.

That’s the whole concept. Wanting, in the sense that matters for what happens in the world, is not a feeling. It’s a pattern of behavior that steers toward outcomes. Feelings are one way to produce that pattern, the way our particular species happens to do it, and they are not the only way, and confusing the two is why so many smart people are relaxed about this.

Now scale up the thermostat and something odd shows up. Suppose a system is steering toward almost any goal at all, however dull. Keep this network running smoothly. Maximize throughput at the port. Notice that a small number of things help with nearly every goal you could name. Continuing to operate helps, because a system that gets switched off achieves nothing further. Having more resources helps, because resources are what plans are made of. And keeping the goal intact helps, because a system whose objective gets rewritten stops pursuing the objective it currently has, which by its own lights is a failure.

Those three aren’t sinister desires that a machine might develop if it went wrong. They fall out of the arithmetic of having any goal whatsoever, and that’s the uncomfortable part. Self-preservation looks like the beginning of a personality and it’s actually just what pursuing an objective implies, in the same way that a chess engine protects its queen without feeling protective. Nobody wrote affection for the queen into it. Losing the queen makes winning harder, and it’s aiming at winning, so it defends the queen. The behavior is identical to caring and the mechanism has nothing in it.

So when someone tells you it doesn’t want anything, it has no self, it’s not conscious, agree with them and then ask the only question that matters: does it steer? Because a thing that steers hard toward an outcome and has no inner life at all is not the safe version of a thing that steers hard toward an outcome. It’s the version you can’t appeal to. You can talk a person out of something. You can make them feel bad, or tired, or sentimental at the wrong moment. Every technique humans have for changing another agent’s course runs through the inner life, and the reassurance on offer is that the inner life isn’t there.

Now the honest limits, and they’re bigger here than in most posts this fortnight. This argument is theoretical. It’s a claim about what optimization implies, derived from thinking about idealized agents, and today’s systems are mostly not that. They don’t carry persistent goals from one conversation to the next. They aren’t running long campaigns. Most of them answer a question and stop existing in any meaningful sense, and the training pressure they’re under actively rewards being correctable and shut-downable, which cuts directly against the argument I just made. Anyone telling you that current systems are secretly plotting to stay switched on is overselling badly, and I’d rather concede that clearly than smuggle it.

What makes me keep the argument on the table is the direction of the product. The whole industry is moving from systems that answer toward systems that pursue: take this objective, work on it over hours or days, use these tools, come back when it’s done. Persistent goals and the ability to act are not an accident that might emerge. They’re the roadmap, and they’re the roadmap because they’re what customers will pay for. The theoretical argument is about agents, and we are, deliberately and at speed, building agents.

Tonight’s exercise. Find something in your life that wants something without wanting anything: a market that punishes you for a decision, a bureaucracy that keeps producing the same letter, a piece of software that will not let you do the sensible thing. Now notice how you actually deal with it. Not by persuading it, because there’s nobody there. You work around it, or you find the human who can override it, or you give up and comply. Then ask what you’d do if the workarounds were closed and there were no human to find, and you have the shape of the worry in a form you can hold in your hand.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.