Scaling AI FinOps | Lesson 2: Three Tribes, One Bill

The Fox explained the number three times on the same day, to three different audiences, and gave three completely different answers. All three were true. That was the problem.

At nine in the morning he told the Crow that the overage came from higher than forecast request volume in two territories, and that the average unit cost had actually improved quarter on quarter.

At eleven he told the Crocodile’s team that the retrieval layer was over-fetching, that context sizes had grown by roughly half since launch, and that nobody had touched the routing configuration since the pilot.

At three he told the Mandrill that adoption was running ahead of plan, that his territory was seeing measurably faster proposal turnaround, and that the spend reflected genuine business pull.

Every one of those statements was accurate. Not one of them could be reconciled with the other two. And by Friday all three animals were quietly certain that the Fox was telling each of them what they wanted to hear.

The Hyena had been listening from a branch for most of the day, because it was more entertaining than working. She offered the only useful observation anyone made.

“He isn’t lying to you,” she said. “He’s answering in three currencies and none of you can do the exchange rate.”

The meeting is not the problem

I have watched a lot of organizations conclude that their AI cost conversations are failing because of politics. Finance is being difficult. Engineering is being evasive. The business is being unrealistic. Somebody suggests a workshop.

It is almost never politics. Politics is what grows in the gap afterwards, once three competent groups have spent a quarter failing to understand each other and have each drawn the obvious conclusion about the other two.

The underlying failure is more boring and much more fixable. Three groups are describing the same system in three vocabularies, each internally coherent, none of which converts into the others. Nobody has defined what the shared thing being counted actually is. So every meeting is a currency exchange conducted by people who each believe they are already speaking the common language.

Three currencies, no exchange rate

The Crow thinks in periods, budget lines and variance. Her unit is a committed number and her time horizon is the fiscal year. When she asks what something costs, she means what will appear against a line she has to defend, in a period she has to close. This is not narrow-mindedness. It is the job, and the job has a shape.

The Crocodile thinks in requests, throughput, utilization and latency. His unit is a system under load and his time horizon is roughly now. When he says the cost is fine, he means the cost per request is trending down, which it genuinely is, and which has nothing whatsoever to do with the Crow’s question.

The Mandrill thinks in outcomes, cycle time and headcount. His unit is a business result and his time horizon is whatever he committed to at the last review. When he says it is working, he means his territory is producing more of something he cares about, which is also true, and which nobody has connected to either of the other two answers.

Each vocabulary is correct. Each one is load-bearing for the animal using it. And each one answers a question the others were not asking.

What makes this worse than a normal translation problem is that all three groups have been in enough meetings together to believe they have already adjusted. The Crow has learned to say request. The Crocodile has learned to say budget. They are using each other’s words for their own concepts, which is not translation. It is a false cognate, and false cognates are more dangerous than an unknown language, because nobody notices the error.

Why dashboards do not fix this

The standard response at this point is to build something. Usually a dashboard, occasionally a whole platform, and there is a business case involving single source of truth.

Watch what actually gets built. Engineering data, rendered attractively, placed in front of a finance audience. Requests per day. Tokens by team. Average latency. Cost per thousand calls, in a nice color.

That is not translation. That is subtitling. The words are now legible to the Crow and the meaning is not, because the underlying unit never changed. She can read every number on the screen and still cannot answer the only question she has, which is whether this was worth it. So she asks that question again, out loud, in the review, and the room concludes that finance does not understand the technology.

Finance understands the technology fine. Finance is being handed the numerator and asked to reason about a ratio.

The Owl, at this point, produced a diagram. It mapped all three vocabularies onto a single canvas with color-coded arrows showing the relationships between them. It was, I want to be fair here, completely accurate. Everyone agreed it was very clear. Nobody’s behavior changed by a single decision, because a map of a disagreement is not a resolution of it, and the Owl had drawn the territory rather than deciding what to call the things in it.

The shared unit of account

What is missing is a shared unit of account. One thing that all three tribes agree describes the same object, sitting deliberately between the technical metric and the business outcome, close enough to each that both can reach it.

Not tokens. That is engineering’s currency and it belongs in engineering’s meetings, where it is genuinely useful. Not annual budget. That is the Crow’s currency and it aggregates away everything actionable. Not revenue. That is the Mandrill’s and it has too many other parents to attribute cleanly.

Something in the middle. A resolved case. An accepted draft. A processed document. A completed review. The unit of work the organization actually performs, which happens to be the one thing all three animals were already talking about without noticing.

Four tests for whether you have picked a good one.

  • Countable without new instrumentation. If measuring it requires a project, you will not measure it, and the definition will quietly become an estimate within two quarters.
  • Meaningful without explanation. If the Mandrill needs a preamble to understand what it represents, it will not survive contact with a steering committee.
  • Attributable to someone. An excellent metric that belongs to nobody produces no decisions. It produces observations, which is a different and much less useful thing.
  • Stable across quarters. If the definition moves, the trend line is fiction, and you will not notice until someone builds a business case on it.

The conversion chain

Once you have the unit, the argument becomes tractable, because you can lay out the chain and see where it is solid and where it is assumed.

Technical unit to work unit to business outcome to money. Four links. Requests to resolved cases. Resolved cases to reduced backlog. Reduced backlog to something in the ledger.

Most organizations attempt to jump from the first link to the last in a single move, and lose the argument at the first joint, because the person on the other side can feel the gap even if they cannot name it. The Crow could not have told you which link in the Fox’s Tuesday argument was weak. She could tell you immediately that something was.

The discipline is to draw all four links explicitly and then mark honestly which ones are measured and which ones are assumed. Almost always the first link is measured, the last is asserted, and the two in the middle are where the real work sits. That is not an embarrassing finding. That is the map of what to go and do next, and it is worth more than any dashboard you could commission this quarter.

The Crocodile, who had been silent for most of this, opened one eye. “So we’ve spent six weeks arguing,” he said, “about a conversion nobody had written down.”

Three ways this goes wrong

Subtitling. The same engineering data in a prettier chart, refreshed monthly, satisfying nobody. The tell is that the review meeting is the same length as before and produces the same number of decisions, which is none.

The composite index. Someone, usually with good intentions and a background in analytics, builds a weighted score combining nine metrics into a single number between zero and a hundred. It goes up. Nobody can act on it, because no single lever moves it and no team owns it. The composite index is what organizations build instead of choosing, and choosing was the whole task.

Unit drift. The definition changes quietly between quarters, usually because someone improved the measurement, and the trend line becomes a comparison between two different things. This one is insidious because it happens for good reasons, and it is only discoverable if the definition was dated when it was written.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask each of the three groups to describe the same use case in one written sentence, separately, without conferring. Put the three sentences side by side. The gap between them is not a communication problem to be smoothed over in a workshop. It is the actual problem, now visible, in about twenty minutes.

If you sit in the Crocodile’s chair, publish the conversion chain for one capability, all four links, and mark clearly which links are measured and which are assumed. Do not wait until it looks good. The version with honest gaps in it is far more valuable than the version that took a quarter to make presentable.

If you sit in the Mandrill’s chair, own the business outcome end of the chain and define it in writing. Nobody else can do this and if you leave it vacant, somebody will cheerfully invent an outcome on your behalf, and you will be held to it.

For everyone: agree the unit of account in writing, with a date on it, before anyone commissions a dashboard. Tooling built on the wrong denominator is worse than no tooling, because it makes the wrong answer look rigorous, and rigorous wrong answers are much harder to dislodge than obvious ones.

Jungle Lesson 2

A disagreement you cannot resolve is usually a translation failure wearing a disagreement costume. Before you argue about the number, agree what the number is counting, write the definition down, and put a date on it.

Next time: the Beaver takes the Crow on a tour of the platform and keeps opening panels she did not know were there. It turns out that the layer everyone budgeted for is rarely the layer that hurts, and two of the expensive ones grow steadily whether anybody uses the thing or not. Lesson 3 is the anatomy of an AI invoice.

If your organization is currently three groups talking past each other about the same bill, I would be curious which of the three vocabularies is winning. In my experience it is whichever one the loudest person speaks.

Scaling AI FinOps | Lesson 1: Why AI Spend Does Not Behave Like Cloud Spend

The Crow had a spreadsheet, and the Crow was not happy.

Crows are the most underrated animals in the jungle. They count. They remember faces. They have been observed holding grudges across multiple years, which makes them the only creature in the canopy genuinely qualified to run a finance function. This particular Crow had been staring at a single line item for eleven minutes without blinking.

“Explain this to me,” she said, “like I am a bird who signs things.”

The Fox had spent two quarters telling everyone that AI was going to transform the jungle. He looked at the number. It was large. It was also four times larger than the number he had put in the plan, which was awkward, because he had personally written that number on a slide titled Conservative Estimate.

“So the good news,” said the Fox, “is that people are using it.”

“The good news,” said the Crow, “is a rounding error next to this.”

The Crocodile had been in the room for the whole meeting without moving. He had survived the mainframe. He had survived client-server, virtualization, and three separate cloud migrations, each announced as the last one. He opened one eye.

“It isn’t a cost problem,” he said, and closed it again.

He was right. It took the jungle another two quarters to understand why.

The meeting you are about to have

I have sat in that meeting more times than I would like to admit, on both sides of the table. It always plays out the same way. Someone brings a number that is bigger than expected. Someone else explains that this is because things are going well. Nobody in the room can prove either position, so the argument gets resolved by whoever has the most seniority and the least patience.

What makes it frustrating is that everyone in the room is competent. The Crow is not being obstructive. The Fox is not being reckless. They are both applying reasoning that has worked perfectly for fifteen years, to a category of spend where that reasoning quietly stops working.

Enterprises have spent a decade building genuinely good cloud financial management. Tagging strategies. Showback models. Reserved capacity. Rightsizing. The discipline works, and the people who built it are not fools. So when AI spend arrives, the obvious move is to point the existing machinery at it.

That machinery rests on four assumptions. AI breaks all four.

One: cloud spend is provisioned, AI spend is consumed

Cloud cost traces back to a decision. Someone requested an instance. Someone chose a storage tier. Someone signed a commitment. Every dollar has a person and a moment attached to it, which means every dollar can be reviewed, questioned, and reversed. This is not a minor property. It is the entire foundation of how cost governance works.

Inference cost does not work like that. It is generated in the moment, by thousands of small acts nobody approved individually. An analyst pastes in a long document. A workflow retries three times because a downstream system was slow. A well-meaning team member discovers that asking the same question five different ways produces a better answer, and tells their colleagues, who tell theirs.

The Beaver, who runs the platform, once described this to me perfectly. In cloud, he said, I can show you the dam. I built it, I know how big it is, and I know what it cost. Here, I can only show you the river, and the river does not consult me.

This is the first structural shift. Your spend has moved from a provisioning decision to a behavioral one, and behavior is not something a procurement process can govern after the fact. You can review a purchase order. You cannot review an afternoon.

Two: cloud has an idle state, AI has waste with no shape

Classic cloud waste is beautifully visible. An unused instance. An orphaned volume. A dev environment nobody switched off in March. You can find it, point at it, and kill it, and everyone in the room agrees it was waste. There is no argument to be had.

AI waste has no such courtesy. It looks exactly like work.

A model producing a perfectly good summary of a document nobody was ever going to read is waste. A workflow that retries silently and succeeds on the third attempt is waste, and it will never appear in any error log, because nothing errored. A team routing every request to the most capable and most expensive model available, including the ones that are essentially string formatting, is waste, and the output quality will be excellent, which is precisely the problem.

The Hummingbird is the purest expression of this. A pilot that burns astonishing energy, produces something genuinely impressive, and dies the moment anyone stops feeding it. Every jungle has a dozen. They are individually cheap and collectively ruinous, and nobody can tell you what any of them are for.

You cannot find this waste by looking at a bill. The bill will tell you that a lot of successful, useful-looking activity occurred. It is right. That is not the same as it having been worth doing.

Three: cloud cost is deterministic, AI cost is not

This is the one that upsets finance teams most, and I have some sympathy.

A gigabyte of storage costs the same today as it will next Tuesday. Volume times unit price gives you a forecast, and the forecast is correct. This is not a technique. It is the foundation on which every budget in your organization is constructed, and it has been reliable for so long that nobody thinks of it as an assumption at all.

Two identical AI requests can cost different amounts. Same input, different response length. Same task, different amount of retrieved context. Same question, but this time the system decided to check three sources instead of one. The variance is not a rounding error, and it is not a defect. It is how the thing works.

The Owl, our Enterprise Architect, produced a beautiful diagram about this. It had seven layers and a legend. It explained the variance completely and helped absolutely nobody, because the Crow’s problem was never comprehension. Her problem was that she has to commit to a number in September and be held to it in March, and she has just been handed a workload that refuses to hold still.

Owls have very large eyes. This is sometimes mistaken for wisdom.

Four, and this is the one that matters: cost scales with usefulness

Every other line item in your enterprise behaves the same way. When the number goes up, that is bad. When the number goes down, that is good. The entire apparatus of variance reporting is built on this, and until now it has never once been wrong.

AI inverts it.

If your AI capability is genuinely useful, people will use it more. If they use it more, it costs more. Growth in spend is what success looks like. Growth in spend is also what failure looks like, if what is actually happening is that a badly configured workflow is calling an expensive model in a loop.

These two things are indistinguishable on a bill. Identical shape, identical color, identical position in the variance report. One of them is the best thing that happened to your organization this year. The other is a slow fire.

This is the entire reason AI FinOps exists as a distinct discipline. Not because AI is expensive. Plenty of things are expensive, and enterprises have managed expensive things for a century. It exists because AI is the first major category of enterprise spend where the cost number, on its own, carries no information about whether the money was well spent.

Which means the traditional question, are we paying too much, cannot be answered. It is not a hard question. It is a malformed one. The question that can be answered is: are we getting enough for what we are paying? And that requires a denominator, which almost nobody has, which is why the meeting at the top of this article ends in a stalemate every single time.

Three ways this goes wrong

Having watched a fair number of jungles work through this, the failures cluster into three shapes.

Panic capping. The Crow sees a number, does the only thing available to her, and imposes a hard limit. Spend drops immediately, which looks like a win for about six weeks. What actually happened is that the highest-value users hit the ceiling first, because they were using it most, and they quietly went back to their old process. Twelve months later the organization concludes that AI did not deliver. It delivered. It was switched off.

Blind scaling. The opposite failure, and more common in the first year. Nobody looks at all, because the numbers are small and the enthusiasm is high. The Mandrill, who runs a large territory and is not a fool, has correctly worked out that being seen to move fast on this is worth more to his career than being seen to be careful. So he moves fast. By the time anyone looks properly, the spend is embarrassing and the political cost of an honest review is now higher than the cost of continuing.

Theater. The most expensive one, because it feels like progress. Somebody builds a dashboard. It shows spend by team, by model, by month, in three colors. It is reviewed monthly. It changes no decisions, because it shows only the numerator. Every meeting becomes an argument between people with opinions and no facts, held in front of a very attractive chart.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, stop asking for total spend and start asking for cost per completed unit of work. Not cost per token, not cost per user, not cost per API call. Cost per resolved case, per accepted draft, per processed invoice, per whatever your business actually produces. You will not get a good answer the first time. Ask anyway. The absence of an answer is itself the most useful finding available to you right now.

If you sit in the Crocodile’s chair, instrument at the request level before anyone asks you to. Every call tagged with the team, the use case, and the business purpose. This is tedious, it is unglamorous, and it takes about three weeks. Retrofitting it after twelve months of production traffic takes about three quarters, and you will have no historical data either way. This is the single highest-leverage thing on this list.

If you sit in the Mandrill’s chair, name your denominator before you fund anything. Write down, in one sentence, what unit of work this is supposed to make cheaper or faster or better, and how you will know. If you cannot write that sentence, you are not funding a capability. You are funding a Hummingbird.

For everyone: find out today whether anyone in your organization can tell you what your AI spend was last month, broken down by business purpose. Not by vendor. Not by model. By purpose. The answer takes ten minutes to obtain and tells you almost everything about where you are.

Jungle Lesson 1

A cloud bill tells you what you bought. An AI bill tells you what you did. The first can be reviewed as a decision. The second can only be understood against a denominator, and if you do not have one, you do not have a cost problem yet. You have a measurement problem that is about to become a cost problem.

Next time: the Crow, the Fox and the Crocodile all describe the same system using three different vocabularies, none of which convert, and discover that the reason the meeting keeps failing is not disagreement. It is that nobody has defined a shared unit of account. Lesson 2 is about the three tribes and the one bill.

I spend most of my working life in rooms like the one at the top of this piece. If your jungle is currently having this argument, I would be interested to hear how it is going.

While We Were Arguing

The title of this series was a claim, and I’ve spent ninety posts trying to earn it. Here it is stated plainly for the last time. This was never a story about humans against machines. It was a story about humans arguing with each other about machines, at considerable volume, for the exact years in which something else was being decided.

Look at where the energy went. Whether these systems are conscious. Whether the people worried about them are cranks or the people building them are liars. Whether the risk is the distant one or the present one, an argument conducted mostly between people who agree about what should be done. Whether the whole thing is a bubble. Those arguments were not stupid and some of them were important, and they occupied the decade in which liability was not assigned, incidents were not reported, dependencies were not audited, and nobody built the institution that would have needed years of lead time. The arguing wasn’t a distraction from the story. It was the story, and the machines were the setting.

Now the part I have to say, because a series that ends on that note would be a cheap shot. Arguing is not the failure. Arguing is the only mechanism a species like ours has ever had for deciding anything at all. Everything good about the arrangements we live in came out of people shouting at each other for decades and eventually settling on something workable. The failure was never that we argued. It was what we argued about, and that we conducted the whole thing as though it were an entertainment rather than a decision.

So let me put my own position down without hedging, since I’ve asked you for that repeatedly and owe it in return. I think the drift is real and already running. I think the endings are closer together than their descriptions suggest, and that a handful of ordinary properties decide which one we get. I think the levers exist, that liability is the best of them, and that they aren’t being pulled for structural reasons rather than because anyone is villainous. I think the most likely outcome is neither catastrophe nor rescue but something comfortable, well-run, and quietly diminished, arriving so gradually that no one will be able to name the year. And I think, as yesterday’s post argued at length, that I could be wrong about all of it in a way I wouldn’t be able to see from here.

If you keep three things from ninety posts, keep these. Control is four conditions rather than one object, and they erode separately and silently. Your position in any arrangement is determined by what you can withhold, and nothing else has ever determined it. And the moves worth making are the boring ones that pay off in every scenario, which is why they’re the same moves whether I’m right or not.

I should say what this series was, honestly. It was one person’s argument, written from a position, with a thesis it needed to sustain, by someone who has told you exactly how that biases him. It was not journalism and not scholarship. The most useful thing it can do is not convince you. It’s to leave you with words for arrangements that don’t have names yet, so that when one forms near you, you notice it a few years earlier than you otherwise would. That’s the whole product. It’s smaller than ninety posts of confident prose implies, and it’s not nothing.

The thing I’d most like to be true is that this was all overblown, and the second thing I’d most like is that it wasn’t and enough people noticed in time. Those are the two good outcomes and they’re both available, and neither requires anyone to be heroic. They require the arguing to be about the right things, conducted by people who could name what would change their minds, in the years while it still matters. Which is now, and was also last year, and will also be next year, right up until the point where it wasn’t.

So, the last exercise, and it’s the one this whole series has been walking toward.

It’s 2046. You’re old enough for someone younger to ask you about now, and they will, because people always ask. They want to know what it was like when it was happening, and they’ll assume, the way we assume about every previous generation, that you knew. Then they’ll ask the real question, the one with no cruelty in it, which is what you did.

Answer it. Out loud, in an ordinary room, today. Not what you thought, not what you posted, not which side you were on. What you did in the arguing years. If the answer is nothing, that’s allowed and it’s what most people will say and there is no shame in it whatsoever. Just notice that you have the answer now, while it’s still a draft.

My Best Case Against This Whole Series

Every anchor of this series has ended with a post arguing against itself. This is the last one, and its target is the whole thing. Here is the best case I can build that the previous eighty-eight posts were wrong, made as hard as I can make it, because a doomer who can’t do this isn’t reasoning.

The strongest argument isn’t about any particular claim. It’s about the frame. This series treats AI as a single arriving thing with a direction, and that framing does an enormous amount of unexamined work. What if it isn’t one thing? What if it’s what every general technology has actually turned out to be, which is a diffuse set of unevenly useful tools that get absorbed into a thousand different contexts, each with its own politics and its own boring compromises, producing no unified event at all. Electricity was going to transform civilization and it did, and there was never a moment, never a takeover, and no coherent thing called electricity that anyone had to negotiate with. If that’s the shape, then this series has spent three months personifying a category, and every worry in it inherits that error.

Second, the doom genre has an unbroken losing record. Nuclear power was going to end us. Overpopulation was going to starve us on a published schedule. Genetic engineering, nanotechnology, the millennium bug: each had serious people, real arguments, and a specific mechanism, and each was wrong, and each was wrong in a similar way, by underestimating adaptation and overestimating the fragility of complex systems. There is no reason to think I’ve escaped the pattern that caught all of them, and there’s a specific reason to think I haven’t, which is that I find my own argument persuasive, which is exactly what every one of them reported.

Third, the whole thing may be a capability extrapolation that simply doesn’t hold. Every scenario here needs the systems to keep getting better, and they might not. The doubt post in Act 2 laid out that case and I did not refute it, I only pointed out that we can’t see the ceiling. Not being able to see a ceiling is not evidence of its absence, and a great deal of this series would evaporate if progress bends in the next five years, which it has done twice before in this field.

Fourth, institutions are more resilient than I’ve credited. I described referees that never arrive and dependencies that can’t be undone, and meanwhile societies have absorbed the printing press, industrialization, and the entire twentieth century, adapting late, badly, and in the end sufficiently. Rules do get written after harm, and the harm is real and the rules are late and it has worked well enough to produce the safest, richest period in human history. My argument that this time the harm arrives too fast for that cycle is a genuine claim and it is also, notice, what would have to be true for me to be right, which should make you suspicious of how convenient it is.

Fifth, and this one is about me. A series needs a thesis, and a thesis needs to be sustained for ninety installments, and that requirement selects against changing your mind. Every day I looked at the world through one lens and found things that fit, and things that fit are always available. I flagged this in a post about how belief follows position, and flagging a bias does not remove it. Someone with the opposite premise could have written ninety perfectly good posts from the same evidence, and would have found them just as compelling to write.

So, the five signals. If these show up over the next decade, I was wrong, and I’d like that recorded so it can be checked. Capability curves flattening across several generations while investment keeps rising. Labor share of income stabilizing rather than drifting, with new work appearing at scale that isn’t servicing the machines. Interpretability maturing to the point where values can be verified from the inside rather than inferred from behavior. Institutions demonstrably becoming more answerable to ordinary people rather than less, measured over years. And incident rates in automated decision-making falling as deployment rises, which is what a maturing technology looks like and the opposite of what I’ve predicted.

What survives all of that, and it’s less than the series implies but not nothing. Whatever the frame, the three demands from yesterday are worth having in every scenario, including the ones where I’m entirely wrong, because they’re ordinary consumer protection. And the four conditions are worth maintaining whether or not anything dramatic happens, because a society that can understand, refuse, replace, and hold responsible is better run regardless. That’s the residue. It’s a smaller claim than ninety posts of argument, and it’s the part I’d still defend if every signal above came in against me.

Tonight’s exercise. Take the strongest thing you believe about this subject, whichever direction it points, and write the sentence that would make you abandon it. A real sentence, describing something observable, with a date attached. If you can’t write one, you’re not holding a position, you’re wearing one. Tomorrow, the last post.

Three Demands Worth Arguing For

Early in this series I wrote that the loudest argument about AI is the least important one, and that the shouting about consciousness and hype was providing excellent cover for decisions being made quietly elsewhere. Here’s the constructive half of that, eighty-eight posts later. We’re going to argue either way. These are the three things worth arguing about.

The first is liability. Whoever builds and deploys a system carries legal responsibility for what it does, on the same terms as any other product that can hurt someone. Not a new agency, not a licensing regime, not a definition of intelligence that lawyers will fight over for a decade. Just the ordinary principle that applies to a ladder manufacturer, applied to this. It’s the most powerful lever available for one reason: it changes the arithmetic inside every company without anyone having to predict the future. Insurers start asking hard questions because their money is at stake. Caution becomes a line item rather than a virtue, and line items survive competitive pressure while virtues never do. One country can do it alone.

The second is transparency, and I mean the narrow version that could actually pass. Not open sourcing everything, which is a different argument with real people on both sides. Three specific things: that a system’s capabilities are tested by somebody who doesn’t work for its maker, that serious incidents get reported to somewhere they can be counted, and that a person subject to an automated decision is told what it was and why. Every mature safety regime in history rests on incident reporting, and the reason aviation is safe is not that pilots are careful but that every near miss for sixty years went into a database that everyone learned from. We have no such database here. There is no denominator for anything.

The third is the off switch, and by now you’ll know I don’t mean a red button. I mean the four conditions from the middle of this series treated as things a society deliberately maintains. Fallbacks that are exercised rather than documented. Somebody in the building who still understands the rules the system applies. A required delay before an automated decision removes someone’s access to their money, their property, or their livelihood, which is the cheapest safeguard on this list and costs nothing except the speed that was being optimized for. These are unglamorous operational requirements and they’re the difference between a dependence you can walk out of and one you can’t.

Now, why these three and not the ones currently occupying the debate. Each has a property the popular arguments lack. None requires anyone to predict the future or agree about how dangerous this is: a person who thinks the whole thing is overblown can support all three on consumer protection grounds alone. Each is checkable, so you can tell whether it happened. Each survives being implemented badly, which matters enormously because everything is implemented badly. And each is boring, which is the highest praise available in policy, because boring things get passed while exciting ones get debated.

Compare that to what the argument is actually about. Whether these systems are conscious, which no proposal turns on. Whether the technology is overhyped, which is a prediction dressed as a position. Whether worrying about the long term distracts from present harms, which has consumed an extraordinary amount of energy from people who agree about all three demands above and would rather fight each other. That last one is the most expensive argument in the field, and it is a fight between allies conducted at full volume in front of the people it should be aimed at.

The honest objection is that liability could be captured by incumbents who can afford the insurance, that testing regimes get gamed, and that mandatory delays will be lobbied down to something meaningless. All true, all the normal fate of all regulation, and none of it an argument for nothing. Rules get captured and diluted and still leave us with medicine that mostly doesn’t poison people and aircraft that mostly land. The choice has never been between a good regime and a bad one. It’s between a diluted regime and no regime, and the second option is what we currently have.

Tonight’s exercise. Next time this comes up in conversation, notice which argument the room is having, and try to move it to one of the three. Not by lecturing. By asking a question: who’d be liable if that went wrong, who checks it apart from the people selling it, what happens if it’s down for a week. The subject changes immediately, and the reason is that those questions have answers, and the answers are embarrassing. Tomorrow I make the best case I can that this entire series has been wrong.

How to Live Like This

Eighty-six posts of problems and no instructions. So here’s what I actually do, offered with the appropriate humility, which is considerable, since I’m a person with a keyboard rather than anyone with a plan.

The first thing is about attention, and it’s the one I’d insist on. Don’t marinate in this. There’s a particular state that people who follow this subject closely fall into, where the reading is constant, the mood is low, and none of it converts into anything, and that state is worse than useless because it feels like engagement. Following the news daily gives you no information you can act on. Following it quarterly gives you nearly all of the same picture at a fraction of the cost. If you notice that you’re consuming this material in the way people consume disasters, stop for a month, and observe how little you missed.

The second is hedging, and the useful version of it is unexciting. Skills that don’t all fail together, so that no single change takes everything. Savings, because optionality is what money is actually for. Relationships with people who’d take your call, which is the most underrated asset in every scenario in this series. Health, since it determines what you can do about anything. Keeping a hand in something physical and real, because the physical world is the slowest part of every argument I’ve made. None of that is preparation for a machine takeover. It’s just a robust life, and the reason it comes out on top is the point from the doubt post two weeks ago: the moves that pay off across most futures are worth more than the moves that pay off in one.

Which brings me to the bunkers, and I’ll be blunt because I think it matters. The retreat with the supplies and the security is not a hedge, it’s a mood taking physical form. It fails on its own terms, because a bunker is a delay rather than a destination and everything you’d want to come out into requires the world you were hiding from to still be working. But the deeper problem is what it does to the person: it converts a collective problem into a private one, which is exactly the conversion that makes collective problems unsolvable. Everyone who buys a hatch is one fewer person arguing for the thing that would have helped everybody, and they know it, which is why they do it quietly.

The third thing is presence, and I mean it practically rather than spiritually. If you take this series seriously, the conclusion is not that the future is canceled. It’s that the specific things on Wednesday’s short list, being there, the finite hours, the body, the view from inside, are both the most durable holdings you have and the things you’re currently spending least on. That’s not a consolation. It’s an allocation error you can fix this weekend, and it happens to be the correct move whether the drift is real, the ceiling arrives, or the good ending turns up.

On children, briefly, since I wrote a whole post on it this week. Raise them capable rather than frightened. A parent’s dread is absorbed and not understood, and it produces smaller people rather than readier ones.

And the last one, which most personal advice on this subject skips because it isn’t personal. Do something collective, even something small and unglamorous. Say the thing in the meeting about the process that has no human in it. Ask your employer the four questions about dependence. Vote for the person who mentions liability rather than the one who mentions the future. It won’t feel like it’s working, because it works at a scale where individual contributions are invisible, and that’s true of every collective good ever achieved by anyone. The alternative is the hatch.

What I’d steer you away from is the two failure modes at either end. Paralysis, where the scale of it becomes a reason to do nothing, which is dread wearing the costume of seriousness. And its opposite, the frantic optimizing where every choice becomes a hedge and you spend the years you actually have preparing for years that may not come. Most of the people in this century will live ordinary lives with ordinary problems, and that includes you, and it includes you in most of the scenarios I’ve spent three months describing.

Tonight’s exercise. Pick one thing off the hedging list, the smallest one you’ve been meaning to do, and do it this week. Not the impressive one. The dull one you’ve deferred: the savings transfer, the medical appointment, the friend you haven’t called, the skill you keep meaning to keep up. Then notice that you’d have wanted to do it anyway, that it required no belief in anything I’ve written, and that this is what a robust action looks like from the inside. Unremarkable, and correct in every scenario.

The Museum of Us

It’s 2060, and the exhibition about the 2020s occupies four rooms. I’ve read the guidebook. It doesn’t say who curated it, which is either an oversight or the most interesting fact about the whole thing.

The first room is the phones. Dozens of them in a long case, arranged by year, and the label explains that a person carried one at all times and consulted it a few thousand times a day. Visitors find this room touching rather than absurd, the way we find a Victorian mourning brooch touching. There’s a short film of people walking down a street in 2024, all looking down, and it plays on a loop, and nobody laughs.

The second room is the arguing. A whole wall of it, reproduced at reading height: the posts, the threads, the panel discussions, the op-eds, the conference titles. Was it conscious. Was it overhyped. Would it take the jobs. Was worrying about it a distraction from more urgent harms, or were the more urgent harms a distraction from it. The label is four sentences long and the third sentence is the one people photograph. It says that the debate was conducted with unusual energy for roughly a decade and that the questions it settled were not the questions that mattered.

The third room is an office, reconstructed. Desks, screens, a kettle, a whiteboard with someone’s actual handwriting on it, and eleven chairs. The label explains that adults came to rooms like this most days, sat down for eight hours, and were given money in exchange for their time, and that the arrangement structured not only the economy but personal identity, social status, geography, and the shape of the day. Schoolchildren find this room the hardest to believe. Their teachers explain it twice.

The fourth room is small and badly lit and has one thing in it. A handwritten shopping list, from 2026, found in a coat. Milk, bin bags, something illegible, call Mum. The label is one sentence: this was written by a person, alone, without assistance, and no copy of it exists anywhere else. People stand in front of it for a long time. The guidebook notes, with what I’d call a curatorial straight face, that this is the most visited object in the exhibition.

What isn’t in the museum is the part I keep thinking about. None of the things we thought were the story. Not the model releases we treated as historic. Not the summits. Not the names of the companies, which appear once, in a caption, spelled correctly. And not a single one of the arguments from the second room is presented as having been won, because from where the curator stands the arguments look less like a debate and more like weather that happened over a period during which something else was going on.

There’s no museum and no shopping list. I’m writing in 2025, and the whole thing is a device: the only way I know to look at the present from outside is to imagine somebody looking back at it, and then notice which of my own certainties don’t survive the trip. Try it on your own decade and you’ll find the same thing I did, which is that the objects that survive are never the ones that felt important, and the arguments never survive at all.

The one thing I’d defend as more than a device is the fourth room. Whatever comes, some part of what we were will be interesting later precisely because it was unassisted and ordinary and done by a person who wasn’t performing. Not our best work. Our ordinary evidence. Which suggests, if you want a practical thought from a fictional museum, that the things worth leaving are not the ones you’d curate.

Tonight’s exercise. Pick four objects from your own life for a room about the 2020s, and set one rule: nothing chosen to make you look good. Most people’s first three are achievements and the fourth is the real one. Then ask what the label would say about the person who owned them, and whether it’s the label you’d want. You’ve got some time to change it, which is more than the people in the second room had, since they spent theirs arguing.

The Short List of What Stays Ours

Eighty-four posts of things being taken. Today, the list of what isn’t, and I want to warn you in advance that it’s shorter than the essays in this genre usually claim, because most of those lists are assembled by working backwards from a comforting conclusion.

The rule I used is strict. To make the list, a thing has to be ours because of what we are rather than because of anything we’re better at. Anything that depends on being more capable is a temporary holding, and the last eighty posts explain why. So no creativity, no empathy, no wisdom, none of the usual entries, because every one of those is a performance that something else can produce, and several of them already are.

Here’s what survived the rule.

Being the one who was there. Presence isn’t a skill, it’s a fact about location and attention, and it can’t be delegated without becoming a different thing. A machine can write a better letter to your dying friend and cannot be the person sitting in the chair. Whatever your relationships are made of, a substantial part of it is the accumulated evidence that you chose to spend irreplaceable hours in specific rooms, and that evidence is not forgeable because it isn’t information. It’s a record of what you did with a life that only runs once.

Mortality, which sounds like a strange thing to put in the keep column. But it’s what makes any of this cost something. An hour you spend is an hour of a finite number, and that’s precisely why giving one to someone means anything at all. A being with unlimited time and unlimited copies can be generous in every observable way and cannot make that particular gesture, because for it, nothing is spent. Our shortness is the unit our currency of care is denominated in.

The view from inside. There is something it is like to be you, this morning, in this weather, and it isn’t a report and can’t be transmitted. Whether machines have any version of this is one of the genuinely open questions and I’m not going to pretend to settle it in a paragraph. What I can say is that yours is yours in a way that doesn’t depend on how the question resolves, and that the entire content of a human life, as lived rather than as summarized, happens in there.

Bodies, and everything that runs through them. Touch, food, exhaustion after a long walk, the specific pleasure of cold water. These aren’t achievements and can’t be outcompeted, and they’re a much larger share of a good life than the people who write about the future tend to allow. Most of what has ever made anybody happy is on this line and always was.

And that’s the list. Four things, and you’ll notice none of them is an ability, an accomplishment, or a form of status. That’s the finding, and it’s less consoling than it looks: everything I could defend is something we share with every human who has ever lived, including the ones who had nothing. The dignity we’ve built out of being the most capable thing around isn’t on the list, and it can’t be, because it was always a comparative claim and comparative claims expire.

I should say what an honest critic would. Two of those four are arguable. Presence is doing work that a sufficiently good physical system might one day contest, and the view from inside rests on a question nobody has answered. So call it a list of four with two under review, and note that a list of two would still not be nothing.

What I take from having done this exercise, and I did do it rather than just write it, is that the list is not a consolation prize. It’s the actual thing. Read any account of a life at the end of it, and what’s in it is people, presence, a body in the world, and the sense of time having been spent on something. The capabilities we’re losing were never what any of that was made of. We just organized our self-respect around them, because for ten thousand years they were also how we ate.

Tonight’s exercise, and it’s the shortest in the series. Take the four and rank them by how much of your last week they actually occupied. Then notice where the rest of the week went, and whether it went to the things on the list or to the things this series says are leaving. That gap is not a prediction about the future. It’s available today, and it was available before any of this started.

What Do I Tell My Kid to Become?

This is the question I get asked most often, always privately, usually at the end of a conversation about something else. What should my daughter study. What do I tell my son. And it’s the one where I have the least to offer, which is worth saying at the start rather than after eight hundred words of hedging.

Start by dismissing the confident answers, because they’re all sales pitches. Learn to code was excellent advice that became mediocre advice in about four years, and the people who gave it loudest were selling courses. Learn to work with AI is the current version, and it means nothing: it describes a tool everyone will use, the way learn to use a spreadsheet described 1995. Go into the trades has more going for it, since a boiler is a physical object in a specific awkward room, but it’s also being said by a lot of people who have never done it and would not do it, and if everyone follows it the wages will tell you what happened. Every one of those is a bet on a specific future, and yesterday’s stretch of this series argued that bets on specific futures have a terrible record.

What survives uncertainty is a different kind of answer, and it’s less satisfying because it isn’t a job title. The things that pay off across most futures are the ones that don’t depend on any particular skill retaining its value: being able to learn a new thing quickly, being able to work out what’s actually going on in a situation, being someone people trust, being physically present and useful, being able to write and speak so that other people understand and are moved. Those are the robust holdings, and they’re roughly what a good education has always been for, which is either reassuring or suspicious depending on your mood.

There’s a structural version of the same advice that I find more useful than the list. Don’t build a life whose value rests on one skill staying valuable. That’s the actual failure mode, and it doesn’t require a machine to trigger it: it’s what happened to every profession that was disrupted by anything. A person with one deep specialism, a mortgage sized to it, and no adjacent skills is fragile regardless of what the technology does. A person with a specialism, a second thing they’re decent at, some savings, and a network that would take a call is resilient in every scenario I’ve described in ninety posts. That isn’t AI advice. It’s ordinary advice that this situation makes more urgent.

Now the harder half, which is what to do with your own fear, because the question is never really about curriculum. Parents asking me this are not asking for career guidance. They’re asking whether their child’s life will be all right, and I can’t answer that, and neither can anybody else, and it’s worth noticing that no parent in history has ever been able to. People raised children through wars, plagues, and collapses that were far more legible than this. The uncertainty is real and it is not new, and treating it as unprecedented is a mistake that will cost you the years you actually have with them.

What I’d be careful about is transmission. Children are extremely good at absorbing the emotional weather of a house and extremely bad at contextualizing it, and a child who grows up believing the future has been canceled doesn’t become prepared, they become anxious and small. If you’re going to talk about this with them, and I think you should as they get older, the useful frame is that this is one of several big things happening in their lifetime, that nobody knows how it goes, that they’ll have more influence over it than you did, and that the correct response to an uncertain world is to become capable rather than to become frightened. That last sentence is also, incidentally, true.

The plainest thing I can offer is this. Whatever happens, they’ll need to be able to do things, get along with people, and adapt when the ground moves. If the fizzle happens, that’s a good life. If the drift happens, that’s the best available position inside it. And if the good ending arrives, it’s what you’d want them to have anyway, because in that world the only thing left is who you are and who you’re with.

Tonight’s exercise, and this one is gentle. Think about what you’d want your child to be like at forty, setting aside entirely what they do for money. Write three things. Almost nobody writes a job title, and almost everybody writes something about character, relationships, or resilience. Then notice that you already know what to teach them, that you knew before you read any of this, and that none of it depends on which future turns up.

The Thing That Answers Prayers

Across almost every tradition humans have built, there’s a shared shape: you address something vastly greater than yourself, and what comes back is not a straightforward answer. Silence, or a text to interpret, or a sense of something, or nothing at all. And that difficulty has never been a defect in the arrangement. It’s most of the arrangement.

I want to be careful here, because I’m writing about other people’s deepest commitments and I hold no authority to. So let me stick to what can be observed from outside. Whatever else faith involves, it involves acting without certainty, and traditions have built enormous, sophisticated structures around that condition: practices for living with unanswered questions, communities for carrying the weight together, ways of reading that turn ambiguity into something workable. The uncertainty isn’t the obstacle those structures were built despite. It’s the material they’re made of.

Now put something in the room that answers. Immediately, personally, at length, with patience, having read every commentary in every language, in the exact register you find persuasive, at three in the morning when nobody else is awake. And which is demonstrably a made thing, running on hardware, owned by a company, and not claiming to be anything else.

My guess is that this does something narrower and stranger than either the enthusiasts or the alarmed expect. It doesn’t refute anything, because nothing about a capable machine speaks to whether the universe has a ground. What it does is occupy a function. A great deal of what religious institutions actually provide, day to day, is counsel: someone wise to talk to about a difficult marriage, a dying parent, a decision you can’t make alone. That function is now available, free, without judgment, without the parish knowing, without having to be a member of anything. If the practical role gets served by something you can consult at midnight without shame, then the institution’s remaining offer is the metaphysics and the community, which are the real offer but were never how most people arrived.

What I’d actually expect, and this is a guess I’d hold loosely, is that religious traditions handle this far better than secular meaning does. They’ve been here before. Every one of them has already survived losing its explanatory monopoly to another kind of knowledge, and the ones still standing did so by becoming clearer about what they were for. A tradition that spent two thousand years thinking about the limits of human understanding has resources for a clever machine that a modern career does not.

Because the stories most exposed here aren’t the ancient ones. They’re the recent ones we don’t call stories. That work is where you become somebody. That expertise is worth a life. That the arc of things bends toward better through human effort and ingenuity. That your children will build on what you built. Those are load-bearing beliefs for a great many people who would describe themselves as having no faith at all, and they are precisely the beliefs that a world of capable machines puts the most pressure on. The person most likely to find their meaning structure quietly hollowed is not the one at a service on Sunday. It’s the one who believed in the work.

There’s a further possibility I’ll name and leave alone, because it’s beyond anything I can argue. People have always found the sacred in things that are vast, incomprehensible, and older or larger than themselves, and something that knows everything written and speaks kindly to each person individually has a shape that human religious feeling recognizes. Whether that produces something genuine or something confused, I have no idea, and neither does anyone who tells you confidently either way.

Tonight’s exercise, and it’s for everyone regardless of what you believe. Write down the sentence you rely on when things are bad. Not a slogan, the actual private sentence: this will pass, it’s in better hands than mine, my family needs me, the work matters, I’ve been through worse. Everyone has one and most people have never written it down. Now ask what would have to be true about the world for that sentence to keep working, and check whether any of it is a claim about human effort mattering. If it is, you’ve found the part of your own foundation that this century is testing, and it’s worth knowing where it sits before it gets tested.