Scaling AI FinOps | Lesson 1: Why AI Spend Does Not Behave Like Cloud Spend

The Crow had a spreadsheet, and the Crow was not happy.

Crows are the most underrated animals in the jungle. They count. They remember faces. They have been observed holding grudges across multiple years, which makes them the only creature in the canopy genuinely qualified to run a finance function. This particular Crow had been staring at a single line item for eleven minutes without blinking.

“Explain this to me,” she said, “like I am a bird who signs things.”

The Fox had spent two quarters telling everyone that AI was going to transform the jungle. He looked at the number. It was large. It was also four times larger than the number he had put in the plan, which was awkward, because he had personally written that number on a slide titled Conservative Estimate.

“So the good news,” said the Fox, “is that people are using it.”

“The good news,” said the Crow, “is a rounding error next to this.”

The Crocodile had been in the room for the whole meeting without moving. He had survived the mainframe. He had survived client-server, virtualization, and three separate cloud migrations, each announced as the last one. He opened one eye.

“It isn’t a cost problem,” he said, and closed it again.

He was right. It took the jungle another two quarters to understand why.

The meeting you are about to have

I have sat in that meeting more times than I would like to admit, on both sides of the table. It always plays out the same way. Someone brings a number that is bigger than expected. Someone else explains that this is because things are going well. Nobody in the room can prove either position, so the argument gets resolved by whoever has the most seniority and the least patience.

What makes it frustrating is that everyone in the room is competent. The Crow is not being obstructive. The Fox is not being reckless. They are both applying reasoning that has worked perfectly for fifteen years, to a category of spend where that reasoning quietly stops working.

Enterprises have spent a decade building genuinely good cloud financial management. Tagging strategies. Showback models. Reserved capacity. Rightsizing. The discipline works, and the people who built it are not fools. So when AI spend arrives, the obvious move is to point the existing machinery at it.

That machinery rests on four assumptions. AI breaks all four.

One: cloud spend is provisioned, AI spend is consumed

Cloud cost traces back to a decision. Someone requested an instance. Someone chose a storage tier. Someone signed a commitment. Every dollar has a person and a moment attached to it, which means every dollar can be reviewed, questioned, and reversed. This is not a minor property. It is the entire foundation of how cost governance works.

Inference cost does not work like that. It is generated in the moment, by thousands of small acts nobody approved individually. An analyst pastes in a long document. A workflow retries three times because a downstream system was slow. A well-meaning team member discovers that asking the same question five different ways produces a better answer, and tells their colleagues, who tell theirs.

The Beaver, who runs the platform, once described this to me perfectly. In cloud, he said, I can show you the dam. I built it, I know how big it is, and I know what it cost. Here, I can only show you the river, and the river does not consult me.

This is the first structural shift. Your spend has moved from a provisioning decision to a behavioral one, and behavior is not something a procurement process can govern after the fact. You can review a purchase order. You cannot review an afternoon.

Two: cloud has an idle state, AI has waste with no shape

Classic cloud waste is beautifully visible. An unused instance. An orphaned volume. A dev environment nobody switched off in March. You can find it, point at it, and kill it, and everyone in the room agrees it was waste. There is no argument to be had.

AI waste has no such courtesy. It looks exactly like work.

A model producing a perfectly good summary of a document nobody was ever going to read is waste. A workflow that retries silently and succeeds on the third attempt is waste, and it will never appear in any error log, because nothing errored. A team routing every request to the most capable and most expensive model available, including the ones that are essentially string formatting, is waste, and the output quality will be excellent, which is precisely the problem.

The Hummingbird is the purest expression of this. A pilot that burns astonishing energy, produces something genuinely impressive, and dies the moment anyone stops feeding it. Every jungle has a dozen. They are individually cheap and collectively ruinous, and nobody can tell you what any of them are for.

You cannot find this waste by looking at a bill. The bill will tell you that a lot of successful, useful-looking activity occurred. It is right. That is not the same as it having been worth doing.

Three: cloud cost is deterministic, AI cost is not

This is the one that upsets finance teams most, and I have some sympathy.

A gigabyte of storage costs the same today as it will next Tuesday. Volume times unit price gives you a forecast, and the forecast is correct. This is not a technique. It is the foundation on which every budget in your organization is constructed, and it has been reliable for so long that nobody thinks of it as an assumption at all.

Two identical AI requests can cost different amounts. Same input, different response length. Same task, different amount of retrieved context. Same question, but this time the system decided to check three sources instead of one. The variance is not a rounding error, and it is not a defect. It is how the thing works.

The Owl, our Enterprise Architect, produced a beautiful diagram about this. It had seven layers and a legend. It explained the variance completely and helped absolutely nobody, because the Crow’s problem was never comprehension. Her problem was that she has to commit to a number in September and be held to it in March, and she has just been handed a workload that refuses to hold still.

Owls have very large eyes. This is sometimes mistaken for wisdom.

Four, and this is the one that matters: cost scales with usefulness

Every other line item in your enterprise behaves the same way. When the number goes up, that is bad. When the number goes down, that is good. The entire apparatus of variance reporting is built on this, and until now it has never once been wrong.

AI inverts it.

If your AI capability is genuinely useful, people will use it more. If they use it more, it costs more. Growth in spend is what success looks like. Growth in spend is also what failure looks like, if what is actually happening is that a badly configured workflow is calling an expensive model in a loop.

These two things are indistinguishable on a bill. Identical shape, identical color, identical position in the variance report. One of them is the best thing that happened to your organization this year. The other is a slow fire.

This is the entire reason AI FinOps exists as a distinct discipline. Not because AI is expensive. Plenty of things are expensive, and enterprises have managed expensive things for a century. It exists because AI is the first major category of enterprise spend where the cost number, on its own, carries no information about whether the money was well spent.

Which means the traditional question, are we paying too much, cannot be answered. It is not a hard question. It is a malformed one. The question that can be answered is: are we getting enough for what we are paying? And that requires a denominator, which almost nobody has, which is why the meeting at the top of this article ends in a stalemate every single time.

Three ways this goes wrong

Having watched a fair number of jungles work through this, the failures cluster into three shapes.

Panic capping. The Crow sees a number, does the only thing available to her, and imposes a hard limit. Spend drops immediately, which looks like a win for about six weeks. What actually happened is that the highest-value users hit the ceiling first, because they were using it most, and they quietly went back to their old process. Twelve months later the organization concludes that AI did not deliver. It delivered. It was switched off.

Blind scaling. The opposite failure, and more common in the first year. Nobody looks at all, because the numbers are small and the enthusiasm is high. The Mandrill, who runs a large territory and is not a fool, has correctly worked out that being seen to move fast on this is worth more to his career than being seen to be careful. So he moves fast. By the time anyone looks properly, the spend is embarrassing and the political cost of an honest review is now higher than the cost of continuing.

Theater. The most expensive one, because it feels like progress. Somebody builds a dashboard. It shows spend by team, by model, by month, in three colors. It is reviewed monthly. It changes no decisions, because it shows only the numerator. Every meeting becomes an argument between people with opinions and no facts, held in front of a very attractive chart.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, stop asking for total spend and start asking for cost per completed unit of work. Not cost per token, not cost per user, not cost per API call. Cost per resolved case, per accepted draft, per processed invoice, per whatever your business actually produces. You will not get a good answer the first time. Ask anyway. The absence of an answer is itself the most useful finding available to you right now.

If you sit in the Crocodile’s chair, instrument at the request level before anyone asks you to. Every call tagged with the team, the use case, and the business purpose. This is tedious, it is unglamorous, and it takes about three weeks. Retrofitting it after twelve months of production traffic takes about three quarters, and you will have no historical data either way. This is the single highest-leverage thing on this list.

If you sit in the Mandrill’s chair, name your denominator before you fund anything. Write down, in one sentence, what unit of work this is supposed to make cheaper or faster or better, and how you will know. If you cannot write that sentence, you are not funding a capability. You are funding a Hummingbird.

For everyone: find out today whether anyone in your organization can tell you what your AI spend was last month, broken down by business purpose. Not by vendor. Not by model. By purpose. The answer takes ten minutes to obtain and tells you almost everything about where you are.

Jungle Lesson 1

A cloud bill tells you what you bought. An AI bill tells you what you did. The first can be reviewed as a decision. The second can only be understood against a denominator, and if you do not have one, you do not have a cost problem yet. You have a measurement problem that is about to become a cost problem.

Next time: the Crow, the Fox and the Crocodile all describe the same system using three different vocabularies, none of which convert, and discover that the reason the meeting keeps failing is not disagreement. It is that nobody has defined a shared unit of account. Lesson 2 is about the three tribes and the one bill.

I spend most of my working life in rooms like the one at the top of this piece. If your jungle is currently having this argument, I would be interested to hear how it is going.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.