Scaling AI FinOps | Lesson 18: Inform, Optimize, Operate

Nothing happened in the jungle for six weeks, and it took the Fox most of that time to notice.

What he eventually noticed was the absence of a particular kind of message. Nobody had asked him for an emergency number. Nobody had called about an unexplained spike. There had been no meeting convened at short notice to work out what something had cost and why.

The weekly review had happened four times and each one had lasted under half an hour. The monthly capability review had produced three decisions, all of them small, all of them made by the people who should be making them. The Crow had reallocated a modest amount of money away from one thing and toward another and nobody had escalated it.

He found this unsettling, which is worth saying plainly, because I think it is the least discussed part of getting this right. Two years of firefighting builds a professional identity around firefighting. When the fires stop, the person who was good at fires has to work out what they are now, and that transition is genuinely uncomfortable for people who have been good at their jobs.

The Crocodile, who had been through this at least twice before under different technology names, was unbothered.

“This is what it looks like,” he said. “You were expecting something more interesting.”

The loop still works, with three changes

The established rhythm of financial operations, inform then optimize then operate, transfers to AI. I want to be clear about that, because there is a fashion for claiming everything is unprecedented and it usually is not.

What changes is not the shape of the loop. It is what goes inside each stage, and in one case, who has to be standing in it.

Inform, adapted

Classic informing tells you cost and allocation. Who spent what, against which line.

That is insufficient here for the reason established in Lesson 1: cost alone carries no information about whether the money was well spent. Informing on AI must carry three things together or it produces the stalemate that opened this series. The cost. The denominator. The acceptance rate.

The reporting object is the unit economic, never the raw spend. If your monthly report leads with total cost, you have built classic cloud reporting and pointed it at a workload where it cannot answer the question anybody is asking.

Optimize, adapted, and this is the big one

Here is the genuine structural difference, and it changes who needs to be in the room.

Classic cloud optimization is largely a procurement and configuration activity. Rightsizing, commitment management, tier selection, eliminating idle resource. Valuable work, and it is mostly done by people who manage contracts and configurations.

AI optimization is almost entirely an engineering activity. Everything in Act III. Routing, context structure, retrieval tuning, caching, agent bounds. Every one of those is a change to how software works, made by people who write software, tested against evaluation infrastructure.

Which means the optimize loop cannot be staffed the way the cloud one was. An AI FinOps function built from procurement and analyst skills will produce excellent reports and will not move the number, because the levers are not in its hands. This is the single most common staffing error in the discipline and it takes about a year to become visible.

Operate, adapted, plus one addition

The governance layer. The value ledger, the reallocation cadence, the guardrail regime, the allocation rule. Everything Act II and the last two lessons built.

And one element classic financial operations does not have, which I would argue is the most important structural decision in this entire lesson.

Cost and quality must be operated together, in the same forum, by the same people.

Because in this domain every cost optimization is a potential quality change. Routing to a cheaper tier might degrade output. Trimming context might lose something that mattered. Reducing agent steps might mean tasks finish less thoroughly. These are not independent variables and they cannot be managed by two groups with separate objectives.

Split them and you get a predictable failure. One group is measured on cost and pushes it down. Another is measured on quality and pushes back. Neither can see the trade, because neither owns both sides of it, and the organization resolves the tension by seniority rather than by analysis. I have watched this consume a year.

The cadence stack

Five rhythms, each with a named owner and a defined decision right. The owner matters more than the frequency.

  • Daily: anomaly detection. Automated. Nobody attends anything. A human is involved only when something fires.
  • Weekly: unit economics per capability. Owned by the capability team. Under thirty minutes if it is working.
  • Monthly: capability review. Cost and quality together. The business attends, which is the difference between a rhythm and a ritual.
  • Quarterly: reallocation, against the ledger, with authority to move money. Lesson 17.
  • Annually: architecture review. Is the placement still right, are the commitments still right, has the crossover moved. Lessons 11 and 20.

Maturity is not linear here

One point that I think is genuinely different and that trips up organizations importing a maturity model wholesale.

Classic maturity models assume you progress. You start immature, you get better, you arrive somewhere. Plan for the destination.

That does not describe an AI estate, because new capabilities keep arriving. A capability that went live last month is at the beginning. One that has been running two years is mature. Both are in your portfolio simultaneously and they need different things.

So the organization is permanently operating at three maturity levels at once, and the operating model has to accommodate that rather than assume a single state. Designing your process for the state of your most mature capability means every new one arrives into a rhythm built for something it is not, and either drowns in governance it does not need or gets waved through because the process assumes a maturity it does not have.

Who runs it

A note that becomes the whole of Lesson 24.

The rhythm works when the capability teams run it and a small central function sets the standards, curates the ledger, and arbitrates. It fails when a central team runs the rhythm on behalf of teams who are not in the room, because then the loop informs people who cannot act and optimizes nothing.

The Fox’s discomfort in the opening of this piece is the good version of this. He was uncomfortable because the work had moved to the people doing it, which is what was supposed to happen, and it left him with less to personally hold. That is the correct outcome and it is not a comfortable one.

Three ways this goes wrong

Cost and quality in separate forums. Each optimized independently, each degrading the other, the trade invisible to everyone and resolved by whoever is more senior.

The cadence nobody owns. A beautiful rhythm on a slide with no name against each loop. Quietly abandoned by month four, and nobody announces it, so it takes another two quarters for anyone to notice.

Single maturity planning. A process designed for the most mature capability in the estate, applied to everything, so new capabilities are governed as though they were established and established ones are governed as though they were new.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, insist cost and quality are reviewed in the same meeting by the same people. This one structural choice prevents more damage than any tool you could buy, and it costs nothing but a calendar change.

If you sit in the Crocodile’s chair, own the optimize loop and staff it with engineers. If your AI FinOps function has no engineering capacity, it is a reporting function and it will not move the number no matter how good the reports get.

If you sit in the Mandrill’s chair, attend the monthly capability review. Attendance by the business is the entire difference between a governance rhythm and a ceremony, and it is visible within two cycles which one you have.

For everyone: put a name against each cadence. Unowned rhythms do not survive a busy quarter, and every quarter is eventually busy.

Jungle Lesson 18

The old loop still works, but optimization has moved from the contract to the code, and cost and quality can no longer be reviewed in separate rooms. If the people cutting the bill are not the people accountable for the output, you have not built a discipline, you have built a tug of war with a budget attached.

Next time: somewhere in the jungle, a well-meaning territory lead imposes a cap, and precisely the thing predicted back in Lesson 1 happens. The best users hit the ceiling first and quietly revert. The Tortoise, who understands ceilings better than anyone because she has spent a career designing them, explains why. Lesson 19 is about guardrails that do not strangle.

The six weeks of nothing happening is the actual goal of this entire discipline, and it is a strange thing to aim for, because there is no way to celebrate it and nobody gets promoted for it.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.