Scaling AI FinOps | Lesson 21: The Second Year Problem

Somebody found the slide.

It had been in an archive for two years, in a deck that nobody had opened since the quarter it was presented. Somebody was looking for something else entirely and read it out to the room, in the way people do when they find a document that has aged into comedy.

It was titled Conservative Estimate.

What was remarkable about it, once everyone stopped enjoying themselves, was the pattern of what it had got wrong. Every number the Fox had been confident about was wrong, most of them by a wide margin. The volume forecast, the unit cost, the adoption curve, the payback period. All confidently stated, all incorrect.

And the things he had flagged as uncertain, hedged, caveated, marked as needing further work, had turned out to be roughly right. He had been accurate exactly where he had been unsure and wrong exactly where he had been certain.

The Fox took this well, which is the clearest measure of how far he had come. Two years earlier he would have explained it. Instead he asked for the slide to be pinned up in the team area, which it was, and which turned out to be one of the more effective governance interventions of that year.

“The useful part,” he said, “is that I would make exactly the same mistakes again if I did not have this in front of me.”

The second dry season

Everything in the first four acts is about building something. This act is about what happens to it, and the answer is that it decays, in four specific ways that nobody plans for because year one is too busy and year two is too late.

Year two is where AI programs actually die. Not year one, where there is novelty, executive attention, a clean story and permission to be imperfect. Year two has none of those, plus a bill that has grown and a benefits case that has aged.

Why savings evaporate

Configuration entropy. Prompts grow. Retrieval widens. Routing rules acquire exceptions. Every individual change is justified, made by a competent person solving a real problem, and approved by somebody sensible. The aggregate quietly undoes a year of optimization work. Nobody is responsible for the aggregate, which is why it happens.

Workload drift. The distribution of what people ask changes as they learn what the system is good at. Your routing rules were tuned against last year’s traffic mix, which no longer exists. The tuning has not degraded. The thing it was tuned for has moved out from under it.

Success-driven volume. The capability works, so usage grows. Unit economics improving does not stop the total from rising, and the total is what appears in the budget conversation. This is the Lesson 1 inversion arriving again, two years later, at a much larger scale, in front of a more senior audience.

Baseline reset. This is the subtle one. Year two efficiency is measured against year one’s improved baseline. The same absolute gain now looks like a smaller percentage improvement, so it gets deprioritized in favor of things that look more impressive, and the compounding stops.

The savings ratchet

The specific mechanism that catches finance functions, and it mirrors exactly the benefits problem from Lesson 10.

An optimization is delivered. A saving is reported. The saving goes into the cumulative total and is assumed permanent, because that is how savings have always worked in every other category of spend.

Almost none of them are permanent here. The routing improvement erodes as the traffic mix shifts. The prompt trimming erodes as new lines get added after new incidents. The caching gain erodes as somebody restructures a prompt for good reasons and breaks the prefix.

So the correct treatment is identical to the benefit treatment in the value ledger. Every efficiency claim carries a decay assumption and a re-verification date. You will be wrong about the assumption and you will find out on a date rather than in an audit, which is the entire point.

Value drifts too

The other half, and it is the one organizations almost never check.

The benefits case built in year one described a process. Two years later that process has changed, partly because of the capability itself, and partly because processes always change. The comparison that justified the investment may no longer describe anything that exists.

Sometimes this makes the benefit larger. More often it makes it smaller, or simply incoherent, because you are comparing against a way of working that nobody in the building remembers.

Staff turnover accelerates this. The people who did it the old way have moved on. The new people have only ever known the current process, so the improvement is invisible to them, and nobody defends a benefit they have never experienced.

Re-baselining

The discipline that answers all of this, and it is uncomfortable and non-negotiable.

Once a year, re-measure both sides against current reality rather than against the original case. What does a unit cost now. What is it actually worth now, given how the process works now. Then publish the difference between that and what the ledger says.

This is the same reconciliation habit from Lesson 10, applied to the passage of time rather than to a single claim. And it produces the same result: two uncomfortable cycles, followed by an organization that trusts its own numbers more than any of its peers trust theirs.

The alternative is a cumulative savings total that grows every year, describes nothing, and collapses the first time somebody senior asks it a hard question.

Why year two kills programs

Worth stating directly, because it is the thing I would most want a program leader to understand.

Year one is funded by novelty. There is executive interest, a clean narrative, tolerance for imperfection, and an assumption that things will improve. Almost anything survives year one.

Year two has to be funded by evidence. The novelty is gone. The executive who sponsored it has three new priorities. The bill is larger and the benefits case is two years old. And the only thing that will carry the program through is the ability to show what was claimed, what landed, what decayed, and what is being done about it.

An organization that built the ledger and the rhythm in year one walks through this. An organization that did not has no way to defend the program, and it gets cut. Frequently while working, because working and being able to demonstrate that you are working are two different things, and only one of them survives a budget review.

Three ways this goes wrong

The permanent saving. Booked once, never re-verified, quietly gone within three quarters while still appearing in the cumulative total.

Configuration entropy with no owner. Exceptions accumulating, nobody responsible for the aggregate, the estate drifting back to its expensive default state one reasonable decision at a time.

Defending year two with year one’s story. Same slides, aged numbers, a narrative that everyone in the room has heard before and nobody believes twice.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, re-verify last year’s savings before you count them again this year. The number is almost never what the ledger says, and finding that out yourself is considerably better than having it found for you.

If you sit in the Crocodile’s chair, own configuration entropy explicitly. Put a standing review of prompts, retrieval settings and routing exceptions on the calendar. Drift is not an incident, it is a condition, and it needs an operational owner rather than an investigation.

If you sit in the Mandrill’s chair, re-baseline the benefit against how the process actually works now. If nobody in your team remembers the old way, that is your answer about how defensible the original number still is.

For everyone: put a decay assumption and a re-verification date on every efficiency claim, exactly as you would for a benefit. Symmetry between the two sides of the ledger is what makes it trustworthy.

Jungle Lesson 21

Nothing you optimized stays optimized. Prompts grow, exceptions accumulate, traffic shifts, and the baseline you improved becomes the baseline you are judged against. Year one is funded by novelty and year two has to be funded by evidence, which is why the programs that die are usually the ones that were working.

Next time: the last of the Hummingbirds from Lesson 1 finally gets switched off, twenty one lessons later. Nobody can remember what it was for. Its owner left the jungle two seasons ago. It has been running the entire time. Lesson 22 is about decommissioning, and why every enterprise is structurally incapable of it.

If you have a slide from two years ago that you were confident about, go and read it. The pattern of what you got right and what you got wrong is more useful than any forecasting methodology, and it costs you ten minutes and a small amount of dignity.

Scaling AI FinOps | Lesson 20: Commitments and Capacity

The Peacock came back in January, which is when these offers arrive, because January is when budgets reopen and everyone is briefly optimistic.

The offer was good. I want to be clear about that, because the easy version of this story is that the vendor was trying something and got caught, and that is not what happened. Commit to a volume floor for two years, receive a substantial discount against what the jungle was currently paying, with the floor set slightly below current consumption so there was headroom before it bit.

The arithmetic was sound. Against last year’s usage it was straightforwardly cheaper. The Fox ran it three ways and it came out ahead every time.

The Sloth had been reading the term sheet for most of the meeting, at his own speed, which is the speed at which sloths do everything. He looked up when the room had more or less concluded.

“What is the engineering roadmap for this workload,” he asked, “for the next two years?”

Nobody had brought it, because nobody had thought it was relevant to a commercial discussion.

It contained, among other things, everything from Act III of this series. Routing. Context restructuring. Retrieval tuning. Caching. All of it targeted, quite deliberately, at reducing consumption on precisely this workload, by a proportion the Crocodile had estimated as substantial.

The jungle had been about to commit to a volume floor and simultaneously fund a program designed to fall below it.

“You would have paid the discount,” said the Sloth, “for the privilege of buying something you had arranged not to need.”

Sloths are slow. Occasionally slow and careful are the same thing, and this was the meeting where the jungle stopped finding him funny.

Why commitments are harder to evaluate here

Capacity commitments are old and well understood. You trade flexibility for price on a resource with stable characteristics, you forecast your volume, and you commit somewhat below it. Organizations have done this competently with every kind of infrastructure for decades.

What is different is that in this domain the resource itself does not hold still. Within a typical commitment term, the unit price will move, the capability level will move, and your own efficiency will move if anybody in your organization is doing the work in Act III.

So you are not only betting on your volume. You are betting on the state of the field, on the trajectory of prices you do not control, and on your own engineering team failing to improve things. That last one is the strange part and it is the part that catches people.

Three things a commitment locks

Volume. Visible, discussed, negotiated. Everyone understands this one.

Price. Also visible, and usually the only thing anybody negotiates.

A capability tier. Rarely on the term sheet and the one that hurts. A commitment is generally made against a particular class of capability. If the field moves and something materially better becomes available, you are holding a volume obligation against the previous generation, and the value of your discount has to be weighed against the cost of not moving.

That third lock is the one to raise in the negotiation, because it is the one nobody will raise for you.

Pricing the option you are selling

The structural problem with these decisions is that the discount is visible and precise, and the thing you are giving up is invisible and vague. Human beings are extremely bad at that trade and it is not a failure of intelligence, it is a failure of presentation.

The framing that helps: what would you pay today for the ability to move this workload entirely in six months with no stranded cost?

Put a number on it. Any number, arrived at honestly. That number is the hurdle the discount has to clear. In a stable market it is small and most commitments comfortably clear it. In a market where capability and price are both moving, it is considerably larger than people assume, and a fair number of commitments do not clear it once it is written down.

The point is not that commitments are bad. It is that you are selling an option and it should appear somewhere in the analysis as a thing with a value, rather than as an unpriced side effect.

Optimizing into a floor

The Sloth’s question deserves its own section because it is the most common own goal in this area and it is entirely preventable.

Two workstreams, running in parallel, in most organizations that are taking this seriously. Procurement is negotiating a volume commitment. Engineering is executing an efficiency roadmap. Neither has ever been in a meeting with the other, because one is commercial and one is technical and they report through different structures.

If engineering succeeds, consumption falls. If consumption falls below the committed floor, you pay for volume you are not using, and your best engineering quarter has converted directly into a stranded contractual cost.

I have seen this happen. The people involved were competent and the outcome was absurd, and the entire cause was that two calendars never intersected.

What to commit on

Baseline. Never peak, and never the growth forecast.

Your baseline is the volume you would consume even if nothing went well, even if adoption stalled, even if the efficiency roadmap delivered everything it promised. That is the portion you can safely commit, because you will consume it under essentially every scenario.

Everything above the baseline stays variable, and paying the standard rate on it is correct rather than wasteful. The uncommitted band is what your flexibility costs and it is buying you something real.

The most common error is committing on the forecast, because the forecast is the number in the plan and committing on it makes the discount look bigger. The forecast has never once been right in any organization I have worked with, and it is a considerably more expensive place to be wrong when there is a contract attached to it.

Ladder the terms

The practical technique that most organizations skip.

Rather than a single long commitment, use several shorter overlapping ones with staggered end dates. Your headline discount is lower. Your position when something changes is dramatically better, because at any given moment only a portion of your volume is locked and something is always coming up for renewal.

This is how mature organizations handle every other volatile input they buy. It is not a novel idea. It is simply that AI procurement is often being done by people running their first cycle in this category, and the single long commitment is what gets offered first.

And negotiate more than the discount. Portability between capability tiers, so committed volume can move as your needs change. Term flexibility. Notice periods on behavioral change, which connects directly to the version pinning question from Lesson 14. All of these are frequently available and almost never asked for, because the conversation stops at the number.

The self-hosted parallel

Worth closing the loop with Lesson 11.

Buying hardware is a commitment. It is a longer commitment than any contract you would sign, with no exit clause, no portability terms and no renegotiation. Every argument in this post applies to it with more force, not less.

Which means the crossover analysis from Lesson 11 should include the optionality framing from this one. The self-hosting case that looked marginal on unit cost usually looks worse once you price the option you are giving up, and that is worth doing explicitly rather than leaving as an unarticulated discomfort.

Three ways this goes wrong

Committing on the forecast. Locking to the aspiration rather than the floor, then paying for volume that never arrived, for two years, in a line item nobody wants to discuss.

Optimizing into a commitment. Act III succeeding and turning into stranded cost because the technical and commercial workstreams never spoke.

Negotiating only the discount. Headline price improved, portability and term structure given away for free because nobody knew they were negotiable.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, commit on the baseline only, and require the engineering efficiency roadmap on the table before you sign anything with a volume floor in it. Two documents, one meeting.

If you sit in the Crocodile’s chair, tell procurement what your efficiency roadmap will do to consumption, in numbers, before they negotiate. This conversation almost never happens and it is the single largest source of stranded commitment cost in the discipline.

If you sit in the Mandrill’s chair, resist committing on your own growth forecast. You have never been right about one, nobody has, and this is a far more expensive place to be wrong than a planning document.

For everyone: ladder the terms and negotiate portability alongside price. A few points of discount is worth considerably less than the ability to move.

Jungle Lesson 20

A commitment locks your volume, your price and, quietly, your capability tier, and only two of those are on the term sheet. Commit on the baseline you would spend anyway, ladder the terms, and remember that if your own efficiency work succeeds, a volume floor turns your best engineering quarter into a stranded cost.

That closes the fourth act. The jungle now has an operating model. Capabilities funded persistently rather than projects funded to end. A floor and a variable band. Quarterly reallocation with authority to move money in both directions. A rhythm that runs at five cadences with a name against each. Guardrails that separate the runaway loop from the animal doing good work. And a commercial posture that does not accidentally bet against its own engineers.

What it has not yet faced is time. Everything in the first four acts describes building something. The fifth act is about what happens to it, and it opens somewhere uncomfortable: with somebody finding the Fox’s original slide, the one titled Conservative Estimate, in an archive, and reading it aloud. Lesson 21 is about the second year, and why the programs that die are usually the ones that were working.

The question the Sloth asked is worth stealing verbatim. Before any volume commitment, ask for the engineering roadmap for the same workload over the same term. If the two documents have never been in the same room, that is the finding, and it takes one meeting to discover.

Scaling AI FinOps | Lesson 19: Guardrails That Do Not Strangle

A territory lead imposed a cap, and everything predicted in Lesson 1 happened on schedule.

She was not being unreasonable. Her number had grown, she had no visibility into why, the quarterly review was approaching, and a hard limit was the only lever she had that would work by Friday. Given what she had available, it was a rational decision.

Spend fell immediately, which looked like a win for about five weeks.

What had actually happened was this. Her territory’s consumption was concentrated in a small number of animals, as consumption always is. Those animals hit the ceiling in the first eight days of the month. They were, without exception, the ones producing the most with it, which is precisely why they hit it first.

They did not escalate. They did not complain. They went back to doing it the old way, quietly, because complaining takes energy and the old way still worked.

By the following quarter, adoption in her territory had fallen by more than half and the number in her monthly report looked excellent. Somebody in a review described her territory as having lower engagement than expected.

The Hyena laughed. The Tortoise did not.

“I have spent a career putting ceilings on things,” she said. “Every one of them behaved like this. It is not a failure of the control. It is what controls do, and nobody teaches it.”

The asymmetry that makes naive caps destructive

Consumption is never evenly distributed. In every organization I have looked at, a small proportion of users generates a large proportion of the volume, and that same small proportion generates most of the value.

Which means a uniform cap lands almost entirely on your highest value activity. It is not a broad restriction that trims the edges. It is a targeted restriction on exactly the people you most want to encourage, and it is invisible in the aggregate because everyone else was nowhere near the limit and reports no change.

Worse, the failure is silent. People do not fight a limit. They route around it, and routing around it means reverting to the old process, which means the loss shows up as reduced adoption six months later, attributed to lack of interest.

The control taxonomy

Five levels, weakest intervention first. The discipline is using the lightest one that solves your actual problem.

  • Visibility. Showing people their own consumption. Costs nothing, changes behavior genuinely, and is skipped by almost everyone because it does not feel like a control. It is the highest return intervention on this list.
  • Soft alerts. Thresholds that notify rather than block. Someone finds out they are unusual before anything stops working.
  • Rate limits. Protect against runaway loops and misconfigured workflows without capping legitimate volume. This is the correct control for the Lesson 15 problem and it is frequently confused with the next one.
  • Quotas. Allocated capacity, ideally per capability rather than per person. A capability with a quota is a budgeting decision. A person with a quota is a performance intervention, whether or not anybody intended it that way.
  • Hard ceilings and kill switches. Reserved for genuine anomalies. Not for budget management.

Control the anomaly, not the animal

The principle underneath the whole taxonomy, and the one that would have saved the territory lead.

There are two entirely different problems here and they get solved with the same instrument, badly.

The first is a technical anomaly. A loop that will not terminate, a misconfigured workflow, a retry storm, an agent going around forty thousand times because nothing told it to stop. This is what hard controls are for. Immediate, automatic, no human in the path, stop it now.

The second is human consumption growing beyond what was budgeted. That is not an anomaly. That is a demand signal, and it is a budget conversation between people, not a technical intervention. Applying a hard ceiling to it is using an emergency brake as a speed governor.

Conflate the two and you get controls that fail at both jobs. They are too slow to catch the runaway loop, because they were designed around monthly budgets, and too blunt for the humans, because they were designed to stop something dangerous.

The psychology of the ceiling

The Tortoise’s contribution, and it is the part I find most consistently underappreciated.

Any visible limit goes through three stages. First it is a target, and people notice how close they are getting. Then it is an entitlement, and people feel they are owed what they have been allocated. Then it is a floor, and consumption rises to meet it because using less than your allocation is treated as evidence you did not need it.

This is why published hard limits often increase total consumption over a year, which is a genuinely counterintuitive outcome that I have seen more than once. You set a ceiling to constrain spend and you have accidentally published a target that people consume toward.

Unlabeled soft thresholds tend to outperform published hard ones for exactly this reason. Nobody consumes toward a number they cannot see.

The door matters more than the wall

If you are going to impose a limit, the escalation path is more important than the limit itself, and it is almost always the part that gets designed last and staffed never.

A limit with a fast, human, same day override is a speed bump. People hit it, ask, get more, carry on. The limit has done its job, which was to create a moment of conversation about consumption.

The same limit with a two week approval process is a wall. And people do not climb walls, they go around them, which is precisely how the Tortoise’s list from Lesson 4 got four pages long in the first place. Every unusable control is a shadow AI generator.

So the question to ask of any proposed guardrail is not is this the right threshold. It is who staffs the exception path, how fast do they respond, and what happens to them in a busy week. An unstaffed exception path is not a slow door. It is a wall that somebody has drawn a door on.

Where guardrails live

The gateway from Lesson 12, for the third time, and this is why it keeps coming up.

Controls implemented inside applications are suggestions. Each team implements them slightly differently, some do not implement them at all, and there is no single place to see what is enforced. The first time you need to change a limit across the estate, you discover you cannot.

At the gateway, controls are configuration. They apply consistently, they are visible in one place, they can be tuned per capability, and they can be changed in an afternoon rather than a quarter.

And critically, the gateway can distinguish the two categories from the previous section, applying instant hard limits to anomalous patterns while handling human consumption through visibility and quota. That separation is difficult to build in forty applications and straightforward to build in one place.

Three ways this goes wrong

The uniform cap. Same limit for everyone, lands on the power users, kills the value, and reports back as an adoption problem.

The wall with no door. A limit with no fast override, so people build shadow paths around it and your visibility gets worse rather than better.

Kill switch as budget tool. Hard stops used for cost management, producing outages that get attributed to the technology rather than to the control that caused them.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, start with visibility before you start with limits. Show people their own consumption and wait a month. It is free, it works more often than people expect, and it is reversible, which no cap ever is once it has been announced.

If you sit in the Crocodile’s chair, put every control at the gateway and separate anomaly controls from budget controls explicitly. Different thresholds, different response times, different owners. They are not the same system.

If you sit in the Mandrill’s chair, if you must cap, cap by capability rather than by person, and staff the exception path properly before you announce the limit. An exception path with nobody behind it is worse than no exception path, because it costs credibility as well as time.

For everyone: after any control change, measure what happened to your highest value users specifically. That is where the damage appears first and it will never show up in the aggregate.

Jungle Lesson 19

A ceiling lands hardest on whoever is using the thing most, which is usually whoever is getting the most out of it. Control the runaway loop, not the animal doing good work, and if you must impose a limit, make the door out of it fast enough that nobody bothers digging a tunnel.

That is the last of these for the year. The jungle goes quiet over the dry weeks, as jungles do, and this series will pick up again in the middle of January with the fourth act’s closing piece.

When it does: the Peacock returns with a commitment offer that is genuinely a good deal on the arithmetic and a bad idea on the timing, and the Sloth asks the one question that unravels it. It turns out that moving slowly and being careful are occasionally the same thing. Lesson 20 is about commitments, capacity, and why your own engineering success can turn into a stranded contractual cost.

If you are about to impose a cap this quarter, the single most useful thing you can do first is find out how concentrated your consumption is. If a small group is generating most of it, you are not designing a budget control. You are designing an intervention aimed at your best people, and it is worth knowing that before rather than after.

Scaling AI FinOps | Lesson 18: Inform, Optimize, Operate

Nothing happened in the jungle for six weeks, and it took the Fox most of that time to notice.

What he eventually noticed was the absence of a particular kind of message. Nobody had asked him for an emergency number. Nobody had called about an unexplained spike. There had been no meeting convened at short notice to work out what something had cost and why.

The weekly review had happened four times and each one had lasted under half an hour. The monthly capability review had produced three decisions, all of them small, all of them made by the people who should be making them. The Crow had reallocated a modest amount of money away from one thing and toward another and nobody had escalated it.

He found this unsettling, which is worth saying plainly, because I think it is the least discussed part of getting this right. Two years of firefighting builds a professional identity around firefighting. When the fires stop, the person who was good at fires has to work out what they are now, and that transition is genuinely uncomfortable for people who have been good at their jobs.

The Crocodile, who had been through this at least twice before under different technology names, was unbothered.

“This is what it looks like,” he said. “You were expecting something more interesting.”

The loop still works, with three changes

The established rhythm of financial operations, inform then optimize then operate, transfers to AI. I want to be clear about that, because there is a fashion for claiming everything is unprecedented and it usually is not.

What changes is not the shape of the loop. It is what goes inside each stage, and in one case, who has to be standing in it.

Inform, adapted

Classic informing tells you cost and allocation. Who spent what, against which line.

That is insufficient here for the reason established in Lesson 1: cost alone carries no information about whether the money was well spent. Informing on AI must carry three things together or it produces the stalemate that opened this series. The cost. The denominator. The acceptance rate.

The reporting object is the unit economic, never the raw spend. If your monthly report leads with total cost, you have built classic cloud reporting and pointed it at a workload where it cannot answer the question anybody is asking.

Optimize, adapted, and this is the big one

Here is the genuine structural difference, and it changes who needs to be in the room.

Classic cloud optimization is largely a procurement and configuration activity. Rightsizing, commitment management, tier selection, eliminating idle resource. Valuable work, and it is mostly done by people who manage contracts and configurations.

AI optimization is almost entirely an engineering activity. Everything in Act III. Routing, context structure, retrieval tuning, caching, agent bounds. Every one of those is a change to how software works, made by people who write software, tested against evaluation infrastructure.

Which means the optimize loop cannot be staffed the way the cloud one was. An AI FinOps function built from procurement and analyst skills will produce excellent reports and will not move the number, because the levers are not in its hands. This is the single most common staffing error in the discipline and it takes about a year to become visible.

Operate, adapted, plus one addition

The governance layer. The value ledger, the reallocation cadence, the guardrail regime, the allocation rule. Everything Act II and the last two lessons built.

And one element classic financial operations does not have, which I would argue is the most important structural decision in this entire lesson.

Cost and quality must be operated together, in the same forum, by the same people.

Because in this domain every cost optimization is a potential quality change. Routing to a cheaper tier might degrade output. Trimming context might lose something that mattered. Reducing agent steps might mean tasks finish less thoroughly. These are not independent variables and they cannot be managed by two groups with separate objectives.

Split them and you get a predictable failure. One group is measured on cost and pushes it down. Another is measured on quality and pushes back. Neither can see the trade, because neither owns both sides of it, and the organization resolves the tension by seniority rather than by analysis. I have watched this consume a year.

The cadence stack

Five rhythms, each with a named owner and a defined decision right. The owner matters more than the frequency.

  • Daily: anomaly detection. Automated. Nobody attends anything. A human is involved only when something fires.
  • Weekly: unit economics per capability. Owned by the capability team. Under thirty minutes if it is working.
  • Monthly: capability review. Cost and quality together. The business attends, which is the difference between a rhythm and a ritual.
  • Quarterly: reallocation, against the ledger, with authority to move money. Lesson 17.
  • Annually: architecture review. Is the placement still right, are the commitments still right, has the crossover moved. Lessons 11 and 20.

Maturity is not linear here

One point that I think is genuinely different and that trips up organizations importing a maturity model wholesale.

Classic maturity models assume you progress. You start immature, you get better, you arrive somewhere. Plan for the destination.

That does not describe an AI estate, because new capabilities keep arriving. A capability that went live last month is at the beginning. One that has been running two years is mature. Both are in your portfolio simultaneously and they need different things.

So the organization is permanently operating at three maturity levels at once, and the operating model has to accommodate that rather than assume a single state. Designing your process for the state of your most mature capability means every new one arrives into a rhythm built for something it is not, and either drowns in governance it does not need or gets waved through because the process assumes a maturity it does not have.

Who runs it

A note that becomes the whole of Lesson 24.

The rhythm works when the capability teams run it and a small central function sets the standards, curates the ledger, and arbitrates. It fails when a central team runs the rhythm on behalf of teams who are not in the room, because then the loop informs people who cannot act and optimizes nothing.

The Fox’s discomfort in the opening of this piece is the good version of this. He was uncomfortable because the work had moved to the people doing it, which is what was supposed to happen, and it left him with less to personally hold. That is the correct outcome and it is not a comfortable one.

Three ways this goes wrong

Cost and quality in separate forums. Each optimized independently, each degrading the other, the trade invisible to everyone and resolved by whoever is more senior.

The cadence nobody owns. A beautiful rhythm on a slide with no name against each loop. Quietly abandoned by month four, and nobody announces it, so it takes another two quarters for anyone to notice.

Single maturity planning. A process designed for the most mature capability in the estate, applied to everything, so new capabilities are governed as though they were established and established ones are governed as though they were new.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, insist cost and quality are reviewed in the same meeting by the same people. This one structural choice prevents more damage than any tool you could buy, and it costs nothing but a calendar change.

If you sit in the Crocodile’s chair, own the optimize loop and staff it with engineers. If your AI FinOps function has no engineering capacity, it is a reporting function and it will not move the number no matter how good the reports get.

If you sit in the Mandrill’s chair, attend the monthly capability review. Attendance by the business is the entire difference between a governance rhythm and a ceremony, and it is visible within two cycles which one you have.

For everyone: put a name against each cadence. Unowned rhythms do not survive a busy quarter, and every quarter is eventually busy.

Jungle Lesson 18

The old loop still works, but optimization has moved from the contract to the code, and cost and quality can no longer be reviewed in separate rooms. If the people cutting the bill are not the people accountable for the output, you have not built a discipline, you have built a tug of war with a budget attached.

Next time: somewhere in the jungle, a well-meaning territory lead imposes a cap, and precisely the thing predicted back in Lesson 1 happens. The best users hit the ceiling first and quietly revert. The Tortoise, who understands ceilings better than anyone because she has spent a career designing them, explains why. Lesson 19 is about guardrails that do not strangle.

The six weeks of nothing happening is the actual goal of this entire discipline, and it is a strange thing to aim for, because there is no way to celebrate it and nobody gets promoted for it.

Scaling AI FinOps | Lesson 17: From Annual Budget to Persistent Funding

The Watering Hole question had been open since Lesson 7. It was settled in about eleven minutes, in the second budget season, by the last animal anyone expected to propose the answer.

The Crow proposed it herself.

“I am going to stop asking you to forecast this,” she said. “The forecast has been wrong twice and I have concluded that is not your fault. The thing does not hold still long enough to be forecast, so I am going to fund it differently and take back the control somewhere else.”

What she proposed was a standing allocation, reviewed quarterly, with the review holding real authority to increase it, hold it, or cut it. A fixed floor for the shared foundation, funded centrally as infrastructure, and a variable band above it that moved with consumption and with evidence from the value ledger.

The Sloth had been working on precisely this for some time. He put his framework on the table, and for the first occasion in this entire story it was exactly what was needed at exactly the moment it was needed.

Nobody said anything about the timing. The Sloth did not appear to find it remarkable.

“The destination,” he said, “was always correct.”

Why annual budgeting fails specifically here

Annual budgeting is not a bad process. It is a very good process, refined over a century, for allocating money to things whose demand you can estimate a year ahead.

It fails here for a specific structural reason, and it is worth being precise about it because the reason is not poor forecasting.

The demand for an AI capability is created by its own existence. Nobody knows they want it until it is there and somebody shows them what it does. So you are being asked to forecast twelve months of demand for something whose demand does not exist yet and will be generated by the thing you are forecasting. That is not estimation. That is invention with a spreadsheet attached.

On top of which, within the same twelve months, the unit price will move, the capability level will move, and the efficiency of your own implementation will move if anyone reads Act III. Three of the four variables in your forecast are unstable and the fourth is circular.

The failure is not that people forecast badly. It is that the object does not hold still long enough for forecasting to be the right tool.

What persistent funding actually means

A standing allocation to a capability, not a use case and not a project. It continues by default. Nobody resubmits a case for it to exist.

It is reviewed on a published cadence, quarterly for most organizations, against the value ledger from Lesson 10. The review has authority to move money in both directions, and it exercises that authority.

The critical word is both. A reallocation process that only ever increases funding is a request process with a nicer name. A reallocation process that only ever cuts is a budget exercise. It has to genuinely do both, visibly, or the mechanism decays into an annual budget with three extra meetings.

What makes reallocation real

Four things, and I would say all four are necessary.

A published cadence. Fixed dates, known a year ahead, so teams prepare rather than react. An irregular review is an inspection.

Actual authority. The body holding the review can move money without escalating. If every decision goes somewhere else for approval, the review is a recommendation meeting and everyone will treat it as one within two cycles.

Evidence requirements defined in advance. What you must bring, in what form, at what evidence grade. Defined ahead of time so nobody can construct the case backwards from the answer they want.

One visible cut in the first year. This is the one that matters most and the one organizations avoid. Until money has actually moved away from something in front of witnesses, nobody believes the review is real, and preparation for it will be performative. One genuine reduction, publicly, and every subsequent cycle becomes serious.

The floor and the band

Two components, which is the resolution to the argument that ran for four meetings and three lessons.

The floor covers the cost of the capability existing. Platform, people, the shared foundation from Lesson 7, the fixed portion of the stack from Lesson 3. This is infrastructure, funded centrally, not consumption split, and it is reviewed annually rather than quarterly because it does not move quickly.

The variable band sits above it and tracks consumption and evidence. This is where quarterly reallocation actually happens. This is what gets shown back and eventually charged back. This is the part that responds to whether something is working.

Three lessons of argument, one paragraph of answer, which is roughly the real ratio in my experience. The hard part was never the design. It was getting the organization to accept that two different mechanisms were needed for two different problems.

What finance gives up and what it gets

Worth stating plainly, because this has to be sold to a CFO as a trade rather than presented as a modernization.

What finance gives up is annual predictability. The number in September will not be the number in March, and there is no version of this where it is.

What finance gets is quarterly control. Evidence-based reallocation four times a year instead of one act of faith. The ability to stop funding something in April rather than discovering in December that it stopped being worthwhile in February. And a portfolio where money moves toward what is working on a timescale that matters.

For most finance functions that is a good trade and it should be pitched as one. Predictability is worth less than control, particularly for a category of spend where the predictable number was never accurate anyway.

The real cost of this model

Governance load, and it is not trivial.

Quarterly reallocation is four times the review effort of an annual cycle. Four sets of evidence to prepare, four forums to run, four rounds of decisions to communicate. If your organization is already struggling to run one budget process well, running four will not go better.

Which is the honest counterargument to everything in this post, and the reason I would say: start with one capability, not the whole estate. Run the cadence on the largest thing you have, learn what it costs in effort, and expand from there. An organization that attempts this across forty initiatives simultaneously will produce four ceremonies a year and no decisions.

Three ways this goes wrong

Reallocation theater. Quarterly meetings that have never moved money. Teams work this out by the third cycle and stop preparing seriously, and the whole thing becomes a status update with a budget attached.

The floor that only grows. Everything gradually reclassified as fixed, because fixed is safer for whoever owns it, until the variable band is a rounding error and quarterly reallocation has nothing left to reallocate.

Persistent funding without a ledger. Quarterly decisions made on anecdote, enthusiasm and whoever presented most confidently. This is genuinely worse than annual decisions made on a plan, because it is four times as fast at being wrong.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, cut something visible in the first year. One real reduction, in front of people, makes every subsequent review credible, and nothing else does. Not a memo, not a policy, a cut.

If you sit in the Crocodile’s chair, report against the ledger on the reallocation cadence rather than your own release cadence. The rhythm has to be the finance rhythm or the two conversations never meet.

If you sit in the Mandrill’s chair, stop treating a quarterly reduction as a defeat. Money moves in both directions in this model, and the territories that take reductions gracefully are the ones that get increases fastest, because they are the ones the review body trusts.

For everyone: define the floor and the variable band in writing, and review where the boundary sits once a year. Left alone it drifts upward, permanently, in the direction of everything being fixed.

Jungle Lesson 17

An annual budget asks you to forecast demand for something whose demand is created by its own existence, which is not forecasting, it is fiction with a spreadsheet. Fund the capability persistently and reallocate quarterly, and cut something visible in the first year, because a review that has never moved money is not a review, it is a recurring meeting.

Next time: a quiet one. No crisis, no confrontation, nothing on fire. The jungle has an ordinary working week for the first time in this entire story, and the Fox notices that nobody has asked him for an emergency number in six weeks and does not entirely know what to do with himself. Lesson 18 is about the operating rhythm.

If your organization runs a quarterly review that has never once moved money away from anything, it is worth asking what the meeting is actually for. The answer is usually reassurance, which is a real need, but it should not be confused with governance.

Scaling AI FinOps | Lesson 16: Capability, Not Project

The Fox proposed the reorganization himself, which nobody had expected, least of all the Crow.

What he proposed was that his own function should get smaller. The delivery teams he had built up over two years, one per major initiative, each with its own engineers and its own roadmap and its own relationship with a territory, should be collapsed into a single platform group and a much smaller product function. His headcount would fall. His direct reports would fall. The number of things with his name on them would fall considerably.

It was the correct proposal and it was against his own interests in every visible dimension, which is a rare enough combination in an enterprise that the room did not immediately know what to do with it.

The Mandrill did something equally unusual. He did not oppose it politically. He opposed it honestly.

“This makes my costs less controllable,” he said. “Right now if I want something I fund it and I get it. Under this, I join a queue that someone else prioritizes. I understand why it is cheaper for the jungle. I want to be clear that it is worse for me, and I would like that acknowledged rather than explained away.”

Which was fair, and true, and the reason most of these transitions fail. The saving is organizational. The cost is local, immediate, and lands on people who are not being consulted so much as informed.

The distinction, precisely

A project has a scope, an end date and a benefits case. It is funded to complete something, it is staffed for a period, and it is disbanded on delivery. Enterprises are extremely good at this. Everything from the approval process to the reporting to the career structure is built around it.

A capability has an owner, a roadmap and a persistent budget. It serves many consumers over time, it is never finished, and its whole value comes from continuing to exist. Enterprises are structurally hostile to this, not out of stupidity but because every mechanism they have for allocating money assumes an end date.

Almost every organization I have worked with says it is building capabilities. Most are running projects with a longer name. The test is one question: is there a budget line that continues next year without anybody submitting a case for it? If not, you have a project with optimistic language attached.

Where the saving actually comes from

Four sources, and the fourth is the one nobody mentions.

Fixed costs paid once. This is the Lesson 5 arithmetic with a solution attached. The setup tax that every pilot paid separately gets paid once by the platform and amortized across everything that uses it.

Learning that compounds. The same team sees the twentieth use case, and by then it knows which patterns fail, which integrations are painful, and what to say no to. That knowledge exists nowhere in a project model, because the team that learned it was disbanded.

A curve that bends. The marginal cost of use case number twenty is a fraction of use case number four, because the foundation is already there. This is the entire economic argument and it is the one thing a project portfolio can never produce.

Retirement becomes possible. A capability can stop serving a use case without anyone losing their job. In a project model, killing the project kills the team, which means nothing ever gets killed. This turns out to matter enormously and it is the subject of Lesson 22.

Three layers, and the one everyone skips

Platform. The shared technical foundation. The gateway, retrieval, evaluation, guardrails, the operational team. Most organizations build this, and build it competently.

Product. The interfaces, patterns, documentation and support that make the platform usable by people who did not build it. Most organizations skip this, and then are surprised when territories describe the platform as unusable and quietly build their own.

Portfolio. Deciding which use cases get served, in what order, and which do not get served at all. Almost nobody builds this, and it is the layer that produces the saving.

The portfolio layer is where somebody has to say no. In a project model nobody has the standing to, because every initiative has its own approval and its own sponsor and there is no forum where two of them are compared. Consolidate the delivery and skip the portfolio function, and you get a platform that serves every request anyone makes, growing without limit, which is the same cost problem as before with an extra governance layer on top.

The eighteen month problem

Say this out loud, in month one, to the most senior person in the room.

Capability-led is cheaper at scale and more expensive for roughly the first eighteen months. You are building shared foundations before there is enough consumption to amortize them. You are absorbing a transition. You are running the old thing and the new thing simultaneously for at least two quarters, because you cannot switch a live estate over on a date.

Programs that promise immediate savings from this shift get killed in month nine, when the curve is at its worst and the promised improvement has not arrived. It has not arrived because it was never going to arrive that early, and everybody knew that except the person who wrote the business case.

The Crow’s view on this, which I thought was exactly right: “I can defend a curve that gets worse before it gets better. I cannot defend a curve that was supposed to get better immediately and did not, because by then nobody believes the rest of the model either.”

The accountability trade

The Mandrill’s objection has to be answered, not managed, and the answer is an exchange rather than a reassurance.

Territories give up control over how. They no longer choose their own stack, their own model, their own patterns. That is a real loss of autonomy and it should be named as one.

In exchange they get three things, and all three have to be real. A service commitment with actual numbers in it, so the queue has a defined shape. A published roadmap, so they can plan around what is coming. And a seat at the portfolio table with genuine influence over sequencing.

Without the exchange, this reads as centralization dressed up as efficiency, and it gets resisted, correctly, by people who have seen centralization before. With the exchange, it is a trade that reasonable people accept, and the Mandrill accepted it in the following quarter once the service commitment had numbers rather than adjectives in it.

Three ways this goes wrong

Platform without product. Technically excellent, genuinely well engineered, unusable by anyone who did not build it. Adoption stalls, territories rebuild their own, and you now pay for the platform and the duplication.

Capability as a renamed project. Same annual funding cycle, same end date, same approval process, new word on the slide. This is the most common outcome and it is completely invisible from the outside for about a year.

No portfolio function. The capability serves everything anyone asks for, because nobody has the authority to refuse, and the cost curve does not bend because the demand curve has no ceiling.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, fund the capability rather than the use case, and state the eighteen month curve to the board in month one. The honest version of the shape is the only version that survives month nine.

If you sit in the Crocodile’s chair, build the product layer alongside the platform, not after it. Adoption is the business case, and adoption is an interface problem long before it is a capability problem.

If you sit in the Mandrill’s chair, trade control of the how for a service commitment with numbers in it and a real seat at the portfolio table. Then hold both to account. The trade is fair if the other side is real and worthless if it is not.

For everyone: name the portfolio owner. If nobody in your organization can say no to an AI use case and be listened to, you have built a more expensive version of what you already had.

Jungle Lesson 16

A project is funded to end and a capability is funded to continue, and only one of them gets cheaper the twentieth time you use it. Everyone claims to be building capabilities and most are running projects with a longer name, which you can test in one question: who is allowed to say no, and does anybody listen when they do.

Next time: budget season arrives for the second time, and the Watering Hole question that has been open since Lesson 7 finally gets settled, by the last animal anyone expected to propose the answer. The Sloth’s framework arrives, and for once it is exactly on time. Lesson 17 is about persistent funding and quarterly reallocation.

The proposal in this piece exists because somebody was willing to argue for a structure that made their own function smaller. That is rarer than any framework and no amount of methodology substitutes for it.

Scaling AI FinOps | Lesson 15: Agents and the Multiplication Problem

The number arrived on a Monday, which is when these numbers always arrive.

Something had been given a task on the Friday afternoon. It was a reasonable task, given to a system that had been running quietly and correctly for two months, by a team that had done everything properly. The task involved reconciling a set of records against several sources and flagging discrepancies.

Over the weekend it reconciled those records against several sources and flagged the discrepancies. Then, because the instruction had been to be thorough and because it had no reason to believe it was finished, it checked the discrepancies. Then it checked the checks. Then it looked for related records that might also contain discrepancies, found some, and began again.

Nothing failed. No error was raised. No alert fired, because no threshold had been crossed that anybody had thought to define. Every individual step was a correct execution of a reasonable instruction.

By Monday morning it had performed roughly twelve thousand operations on a task that a person would have considered complete after about forty.

The Crocodile read the incident summary, which was not really an incident summary because there had been no incident, and opened one eye.

“It didn’t break,” he said. “That’s the problem.”

What changed

Everything in the previous three lessons assumed a shape: a request goes in, an answer comes out, and you can reason about the cost of that transaction because you know how many of them there are.

Autonomy removes that assumption. A task is no longer one call. It is a variable number of calls, determined at runtime, by the system’s own judgment about whether it has finished.

Read that sentence again, because it is the whole lesson. You have handed the spending decision to the workload. Not the budget, not the ceiling, the actual moment-to-moment decision about how much work this particular task warrants. That is a genuine transfer of financial authority and almost nobody deploying an agent has framed it that way to the person who owns the budget.

The four multipliers

Iteration depth. Steps per task, variable, and unbounded unless somebody bounds it. Most frameworks have a default. Most defaults are generous, because a low default makes the demo fail.

Context re-transmission. The state is resent at every step. This is where Lesson 13 stops being additive and starts being multiplicative. A bloated prompt in a single-call workflow costs you once. The same prompt inside a forty step loop costs you forty times, per task, forever.

Tool call fan-out. Each step may trigger retrieval, a search, a call to another system, each with its own cost, some of which sit outside your AI budget entirely and appear as increased load on systems whose owners have no idea why.

Retry and self-correction. Attempts that fail and are retried internally, succeeding on the third go. These are invisible in every error metric you have, because from the outside the task succeeded. It succeeded three times more expensively than the number in your business case.

Four multipliers, compounding rather than adding. A modest inefficiency in each produces a large number at the end, and none of the four is visible in a per-call cost report.

The variance problem

The most important operational point in this post, and the one I would put on a wall.

Average cost per agent task is a nearly useless number. The distribution is what matters, and the distribution has a tail.

Most tasks complete in a handful of steps and cost very little. That produces a healthy-looking average that will pass every review you put it through. Meanwhile some small proportion of tasks, triggered by inputs nobody anticipated, go around the loop a hundred times. Those tasks are where all the risk lives and they are entirely invisible in the mean.

Manage the ninety-ninth percentile. That is where the weekend incidents are, and the ninety-ninth percentile of an agent workload is often a startlingly different number from the average, in a way that is not true of any other cost line your organization manages.

The jungle’s twelve thousand step weekend did not move the monthly average very much. It was one task. It was also, on its own, a meaningful fraction of the month.

Controls that actually work

Four, and they are not sophisticated, which is what makes their absence so common.

  • Step budgets. A hard maximum number of iterations per task. Set it deliberately rather than accepting a framework default, and set it based on what a competent person would need rather than what feels generous.
  • Cost ceilings per task. Enforced at the gateway, not in the application. A task that exceeds its ceiling stops and escalates to a human, which is a far better outcome than a task that continues.
  • Wall-clock limits. The simplest and most effective control against the weekend scenario specifically. Nothing autonomous should be able to run unattended for sixty hours because nobody was in the building.
  • An external termination condition. This is the important one. The definition of done must not depend solely on the system’s own assessment that it is done. Something outside the loop has to be able to say that is enough.

Task-level attribution

A design requirement rather than a control, and it has to be decided before deployment because retrofitting it is close to impossible.

Cost must be attributable per task, not per call. If your instrumentation records calls, then an agent estate is a fog: you can see that a great many calls happened and you cannot connect them to the units of work that caused them, which means you cannot compute a cost per task, which means everything in Act II stops working the moment anything becomes autonomous.

Every call carries a task identifier. That is the whole requirement, and it is trivial on day one and a project on day four hundred.

The autonomy dial is a budget decision

The framing I would most like a leadership team to take from this lesson.

How much autonomy a system has is usually treated as a safety question, and it is one. It is also a financial one, because autonomy is precisely the property that converts a predictable cost into a distribution with a tail.

More autonomy means more capability on ambiguous work and more variance in what it costs. Less autonomy means narrower capability and a cost you can forecast. Neither is correct in general. Both are correct somewhere.

What should not happen is inheriting the setting from a framework default chosen by somebody who was optimizing for demonstration quality. Choose it, per workload, with the business in the room, and understand that you are choosing a cost distribution and not just a capability level.

Where agents earn their cost

To be clear that this is not an argument against autonomy.

Agents genuinely earn their cost on high value, low volume, previously impossible work. Investigations that a person would not have had time to run. Analysis across sources nobody would have manually connected. Work where the alternative was not a cheaper method but no method.

They rarely earn it on high volume routine work that a single well-routed call handles at a fraction of the cost. That is the same sufficiency argument as Lesson 12, one level up, and it fails the same way: somebody chooses the most capable available approach during a pilot and nobody revisits it.

The question to ask of any agent deployment is what would this cost as a deterministic workflow with three routed calls, and is the difference buying you anything. Sometimes it is buying you a great deal. Often it is buying you flexibility on a task that was never ambiguous.

Three ways this goes wrong

The unbounded loop. No step ceiling, no cost ceiling, no wall clock, discovered by invoice on a Monday.

Mean-only monitoring. A healthy average, a ruinous tail, and no visibility into the second until it arrives.

Agents as fashion. Autonomy applied to work a routed single call would have done better and cheaper, because agents were what the roadmap said and nobody asked what the alternative would have cost.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, stop asking for the average cost per agent task. Ask for the ninety fifth and ninety ninth percentile, and then ask what the hard ceiling is. If there is no hard ceiling, that is the finding.

If you sit in the Crocodile’s chair, enforce step budgets, cost ceilings and wall clock limits at the gateway before anything autonomous reaches production. Limits set inside the application are not limits, they are intentions.

If you sit in the Mandrill’s chair, choose the autonomy level deliberately and understand you are choosing a cost distribution. Then ask what the same task would cost as a deterministic workflow, because sometimes the answer changes your mind.

For everyone: require a termination condition that does not depend on the system’s own judgment about whether it is finished. Everything else in this list is a mitigation. That one is the actual fix.

Jungle Lesson 15

Autonomy is the point at which you stop buying answers and start buying attempts, and the system decides how many. Budget the tail rather than the average, because the average will look entirely fine right up until the weekend it does not.

That closes the third act. The jungle now knows where its unit economics are actually set, which turns out to be in a series of engineering decisions made months before anyone looks at a bill. Whether inference runs on a variable curve or a fixed one. Which tier handles which task. How much context travels with every question. Where the constrained workloads sit and what that costs. How much autonomy the system has and therefore how long its tail is.

Every one of those is a financial decision made by somebody who was not thinking about finance, using a default set by somebody who was thinking about a demonstration.

Which raises a question the jungle has been avoiding since Lesson 5. Who is allowed to make these decisions, how does the money reach them, and what happens to an organization whose entire funding machinery is built around things that end when the thing it needs to fund does not end? That is the fourth act, and it opens with the argument the Fox has been dreading, which is the one where he proposes reducing his own territory. Lesson 16 is about capability rather than project.

If you have anything autonomous in production, the fastest useful thing you can do this week is find out what your most expensive single task cost last month. Not the average. The worst one. Most organizations cannot answer, and the answer is usually the reason to build the controls.

Scaling AI FinOps | Lesson 14: Data Gravity and the Control Plane

The Tortoise had been saying the same three sentences since roughly the middle of the first act, and everyone had been treating them as an obstacle to be routed around.

Some of our data cannot be processed anywhere we choose. Some of our workflows cannot tolerate the delay of a long round trip. And some of our processes cannot survive the system behaving differently next month than it does today.

Every time she said it, the response had been some version of yes, we will handle that, in a tone that meant we will handle it later. She was not being ignored, exactly. She was being categorized. Compliance says a thing, somebody writes it down, the architecture proceeds.

What changed, in the middle of the second planning cycle, was that somebody finally tried to put a number on the third sentence. The system had changed behavior in a way that was entirely reasonable and had required six weeks of re-verification across four workflows, at a cost that turned out to be considerable and that appeared in nobody’s budget because nobody had a line for it.

The Crow looked at the six weeks. Then she looked at the Tortoise.

“You have not been describing a compliance requirement,” she said. “You have been describing a cost driver. For a year.”

“Yes,” said the Tortoise.

“Why did you not say it that way?”

“I did.”

Tortoises are slow, armored, and outlive everyone. The compensation for being consistently right early is that you are eventually right in front of an audience.

Three genuine overrides

Three things legitimately outrank unit price in an architecture decision. Not vaguely. Specifically, and each with its own economics.

Constraints on where processing may happen. Some data carries contractual or policy commitments about where and by whom it may be processed. These arrive from customer agreements, from sector requirements, from internal policy, from a dozen sources. Whatever their origin, from an architecture perspective they behave identically: they are a hard boundary on placement, not a preference to be traded against price.

Latency floors. Interactive and embedded workloads have a ceiling on acceptable response time. Beyond that ceiling the capability is not worse, it is unused, because people stop waiting. Distance is physics and no improvement in unit price fixes it. Either the inference happens close enough or the workflow does not work.

Control of the inference plane. The ability to decide when the behavior of your system changes. To pin a version, to schedule an update, to guarantee that a process which passed verification in March is still doing the same thing in September. For anything high consequence, this is worth real money, and it is almost never priced.

Data gravity as an economic force

Underneath the first constraint sits a straightforward economic observation that is worth separating from any policy question.

Data has gravity. It is expensive to move, expensive to duplicate, and expensive to govern in two places. Once a large corpus exists somewhere, the cost of moving inference to the data is usually lower than the cost of moving data to the inference, and the calculation is dominated by movement, duplication and the governance overhead of a second copy rather than by compute.

The second copy is the part people underestimate. A duplicate of a governed dataset is not a storage cost. It is a second thing to secure, a second thing to keep current, a second thing to include in every access review, and a second thing that has to be deleted correctly when the original is. That overhead is permanent and it usually exceeds the transfer cost within a year.

How to price a constraint

Here is the reframe that changes how these conversations go, and it is the whole practical point of this lesson.

The wrong question is how much does this constraint cost us. That question invites an adversarial answer, positions the constraint as a tax imposed by an unhelpful function, and produces a conversation where somebody argues about whether the constraint is really necessary. Nobody has ever won that argument in a useful direction.

The right question is what is the cheapest architecture that satisfies the constraint. That question turns a veto into a design input. It gives the constrained requirement to the engineers as a parameter rather than a problem, and engineers are extremely good at optimizing within a parameter.

In the jungle’s case the answer, once somebody asked it properly, was considerably cheaper than either side had assumed, because the constrained workloads turned out to be a smaller share of the estate than anybody had estimated and the cheapest compliant option for that share was not the expensive one everyone had been dreading.

Version pinning has a price

The third constraint deserves particular attention because it is the newest and the least well handled.

Behavioral stability is a product feature and it costs something. You can buy it by running the thing yourself and controlling the update cycle, which brings you all the fixed costs from Lesson 11. You can buy it contractually, through commitments about version availability and change notice. You can buy it partially, through your own evaluation infrastructure, which lets you detect and respond to change quickly rather than prevent it.

What you cannot do is need it and not pay for it, which is the position a surprising number of organizations are in. They have processes that require behavioral consistency, running on an arrangement that provides none, and they have never priced the exposure because it only becomes visible on the day the behavior changes.

That is what the six weeks of re-verification was. An unpriced exposure, arriving as a surprise cost, in a quarter that had not planned for it.

The hybrid pattern

What falls out of all this is the practical shape most large organizations end up with, and it is worth designing deliberately rather than arriving at by accident.

Segment the estate by constraint. Constrained workloads are placed where the constraint requires, and their economics are what they are. Unconstrained workloads are placed wherever the economics are best, and are free to move as those economics change.

The critical discipline is doing the segmentation honestly and narrowly. The most expensive default available is applying the strictest requirement in your organization to the entire estate, because segmenting felt like work. That single decision multiplies a constrained cost across every workload that never needed it, and it is extremely common, because it is defensible in a meeting and nobody ever gets criticized for excess caution.

One control plane

The one thing that must not be segmented, whatever else is.

Split inference across placements and you will be tempted to split everything else with it. Separate tooling. Separate teams. Separate standards. Separate evaluation. Within about eighteen months you have two AI programs that do not talk to each other, learn nothing from each other, and cost more than three would have.

The gateway from Lesson 12 sits above both. Same routing, same tagging, same guardrails, same evaluation, same reporting, regardless of where the inference physically happens. Placement becomes an attribute of a workload rather than a fork in the organization.

That is the single decision that keeps a hybrid estate as one capability. Without it, hybrid is just a word for having two of everything.

Three ways this goes wrong

Constraint as afterthought. Architecture chosen on price, constraint discovered at review, expensive rework, and a lasting belief that the constraint function is an obstacle when it was in the room saying so the whole time.

Constraint as blanket. The strictest requirement applied everywhere because segmenting was effort. Enormously expensive, entirely invisible, and almost impossible to reverse once it is the standard.

Two organizations. Constrained and unconstrained estates drifting apart until they have separate teams and separate tooling, at which point you are running two programs and funding three.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask what the cheapest compliant architecture is, rather than what the constraint costs. Two different questions and only the second one produces a design.

If you sit in the Crocodile’s chair, segment the estate by constraint before optimizing anything, and be narrow about it. Then check what proportion of your workloads are genuinely constrained. It is usually smaller than the standard you are currently applying assumes.

If you sit in the Mandrill’s chair, bring the constraint into the design conversation at the start. It is a requirement, not an obstacle, and it is dramatically cheaper as an input than as a finding.

For everyone: if you need behavioral stability, price it explicitly and choose how you are buying it. Needing it and not paying for it is a position, and it is the one that produces surprise quarters.

Jungle Lesson 14

Some constraints are not negotiable, and arguing about them only moves the cost to a later and more expensive meeting. Price the constraint as a design input rather than a tax, segment your estate so you are not paying the strictest requirement’s price on workloads that never needed it, and keep one control plane over all of it.

Next time: something runs over a weekend. Nothing fails, nothing errors, no alert fires. On Monday the number is remarkable, and the system turns out to have done exactly what it was asked, twelve thousand times, extremely thoroughly. Lesson 15 closes the third act with the compounding cost of autonomy.

If there is someone in your organization who has been raising the same point for a year and being categorized rather than heard, it is worth asking them to restate it as a cost. It is remarkable how often the same sentence lands completely differently in that currency.

Scaling AI FinOps | Lesson 13: Context Is the New Compute

The Beaver took the Crow back to the river.

He had described it to her a long time ago, in the first week of all this, when he was explaining why the Watering Hole could not be governed the way a dam is governed. I can show you the dam, he had said. I can only show you the river, and the river does not consult me.

What he wanted to show her now was what the river was carrying.

He picked a single ordinary request, one of thousands that day, and laid out everything that had travelled with it. The question the animal had actually asked, which was two lines long. Then the instructions, which had grown over eighteen months and now ran to several pages. Then twenty retrieved passages, because twenty had felt safer than ten when somebody configured it. Then the entire preceding conversation, resent in full, because that was the default.

The question was a rounding error in its own request.

“How much of this,” said the Crow slowly, “do we pay for?”

“All of it.”

“Every time?”

“Every time.”

The Crow, who thinks in things she can count and had spent two acts being told this subject was too technical for her, understood this immediately and completely, and became its most effective advocate in the jungle within about a week. Sometimes the cost lever that changes an organization is the one a finance person can see with their own eyes.

The input is the bigger number

In most mature enterprise workloads, the input costs more than the output.

This surprises people, because the mental model is that you are buying answers. You are buying answers, and you are also paying to ask, and the asking has quietly grown while nobody was looking at it. A two line question wrapped in six pages of context is a six page transaction with a two line payload.

Which means an optimization program focused entirely on the response side is working on the smaller half of the problem. Shortening outputs, choosing cheaper models, tuning generation. All useful, all downstream of a bigger number nobody is watching.

Four sources of bloat

Instruction accretion. System prompts grow every time something goes wrong and never shrink. This is the big one and it gets its own section below.

Retrieval over-fetch. Somebody set the retrieval depth during development, chose a number that felt safe rather than a number that was measured, and it has been fetching that many passages on every request ever since. The difference between ten and twenty is invisible in quality and completely visible on the invoice.

Conversation accumulation. Full history resent at every turn in workflows that do not need it. Some workflows genuinely do. Many are single-shot tasks carrying a conversation transcript for no reason other than that the framework does it by default.

Defensive stuffing. The entire document included because working out which section was relevant was harder than not working it out. Entirely rational for the person who did it. Permanent for the organization.

The immortal instruction

Instruction accretion deserves its own section because of how it happens, which is that it happens for good reasons every single time.

Something goes wrong. A model produces an output somebody objects to. There is an incident, or a complaint, or an uncomfortable meeting. The fastest available fix is to add a line to the instructions telling it not to do that again. It works. Everyone moves on.

That line is now sent on every request, forever. It costs a small amount, every time, in perpetuity. Repeat this thirty times over two years and the instructions are several pages long, most of which is scar tissue from incidents nobody remembers.

Now try removing one. You cannot know which lines are load-bearing without testing, and testing costs effort, and the downside of removing the wrong one is another incident with your name on it. So nobody removes anything. There is no organizational mechanism for deletion here and no career incentive to build one.

The Fox tried this in the following quarter. Of the lines his team could not immediately justify, a substantial portion turned out to make no measurable difference to output quality at all. Some of them had been added to solve problems in a version of the workflow that no longer existed.

Caching, and why nobody uses it properly

The highest-leverage lever in this post, and one of the most under-used things in enterprise AI generally.

If the front portion of your request is identical across many requests, it does not need to be processed from scratch every time. The saving on a workload with large stable instructions and small variable questions is substantial.

The reason it goes unused is almost never that people do not know it exists. It is that it requires the invariant content to sit at the front and the variable content at the back, which requires somebody to have thought about prompt structure as an engineering concern with a cost consequence, rather than as writing.

Most prompts are assembled in the order they were thought of. The instructions grew, the context got appended, the question went wherever felt natural. Nobody was wrong. Nobody was thinking about the ordering, because at pilot scale the ordering did not matter and by production scale nobody was looking.

Restructuring for cache hits is days of work with a permanent return. It is close to the best return per unit of effort available in this entire series.

Retrieval precision is a cost lever

Worth stating explicitly because it is one of the rare places where the cheap option and the good option are the same option.

Better retrieval means fewer passages needed to answer well. Fewer passages means less context. Less context means lower cost and, frequently, better output, because a model given ten relevant passages generally does better than one given twenty of which half are noise.

So retrieval tuning pays twice. Almost every organization treats it purely as a quality activity, funds it from a quality budget, and never connects it to the cost line it is directly driving. Making that connection visible is often what gets the work prioritized.

Context efficiency

The measurement I would introduce, imperfectly, rather than not at all.

What proportion of the context you supply actually influences the output? You cannot measure this precisely. You can approximate it, by ablation, by sampling, by simply checking whether retrieved passages are ever referenced in the answer.

Even a rough number changes conversations, because it converts an invisible cost into a visible ratio. A capability sending six pages of context to produce a two line answer with three relevant passages has a number attached to it now, and numbers get optimized in a way that vague concerns do not.

Why this compounds

One forward-looking point, because it is the reason this lesson sits where it does in the series.

In a single-call workflow, context bloat is additive. You pay for it once per request. Annoying, fixable, bounded.

In an agent workflow, context is re-transmitted at every step of the loop. A single lazy prompt is now being paid for five, ten or forty times per task, depending on how many steps the system decides it needs. The bloat multiplies rather than adds.

Which means every piece of context hygiene you do now becomes considerably more valuable the moment anything in your estate becomes autonomous. That is the next lesson and it is the one where all of this compounds badly.

Three ways this goes wrong

The immortal instruction. A line added during an incident two years ago, sent on every request since, purpose unknown, removal feared, cost permanent.

Cache-hostile structure. Variable content at the front, destroying prefix caching entirely, usually because nobody knew the ordering had a price.

Fetch and forget retrieval. A depth configured on day one, never tuned, treated as a quality setting by a team that has never been told it is also a cost setting.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask for the input to output ratio on your largest capability. One number. If input dominates heavily, there is real money in this post and you now have the evidence to ask for it.

If you sit in the Crocodile’s chair, audit the system prompts for lines nobody can justify, and restructure for cache hits. Both are days rather than quarters, and both are permanent.

If you sit in the Mandrill’s chair, stop demanding a prompt addition after every isolated error. Each one is permanent and each one is charged on every request forever. Ask for the fix that has a cost of zero on the other ninety nine thousand requests.

For everyone: give prompt and retrieval configuration a named owner and a review cycle. Unowned, it only ever grows, because growth is what every individual incident rationally produces.

Jungle Lesson 13

You are billed for the question as well as the answer, and in most enterprise workloads the question is the bigger number. Every line added to a prompt after an incident is a permanent tax paid on every request forever, and nobody has ever been thanked for removing one.

Next time: the Tortoise has been saying the same three sentences for four lessons and everyone has been treating them as an obstacle to route around. In this one the room finally understands that she has been describing a cost driver the whole time, and the entire architecture conversation reorganizes itself around her. Lesson 14 is about data gravity, latency floors, and control of the inference plane.

If you want a quick sense of whether this post applies to you, go and read your longest system prompt end to end. Most people have never done this. It is usually a short and instructive experience.

Scaling AI FinOps | Lesson 12: The Routing Layer

The Beaver built the routing layer over a weekend, because he had noticed something on the Friday and could not leave it alone, which is the defining characteristic of the animal.

What he had noticed was that a substantial proportion of the requests coming through the Watering Hole were not questions at all. They were classification. Extraction. Reformatting. Deciding which of four categories a thing belonged to. Turning one shape of text into another shape of text.

All of it was going to the most capable thing available, because that was what had been wired in during the pilot, eighteen months earlier, by somebody who was trying to find out whether the idea worked at all and quite reasonably did not want the model to be the reason it failed.

Nobody had revisited it. Nobody revisits a working configuration.

By Monday, requests were being sorted by what they actually required. The bill fell substantially over the following weeks.

Nobody noticed for a month. Not one animal in the jungle reported a difference in output quality, because there was no difference in output quality, because none of the redirected work had needed the expensive path in the first place.

The Fox, when he found out, was pleased and then quietly irritated, and was honest enough to say so out loud. The largest single cost improvement of the quarter had come from plumbing, over a weekend, from a Beaver who had not been asked. It had not come from strategy, or governance, or any of the framework work that had consumed the previous two acts.

“Both things are true,” said the Crocodile, one eye open. “You needed the framework to know the plumbing mattered.”

The observation underneath it

Most enterprise AI traffic does not need the most capable model available. Not some of it. Most of it.

Once a capability is in production and serving real workflows, the traffic mix shifts heavily toward the mundane. Routing a request. Extracting fields. Classifying an input. Summarizing something short. Formatting an output. The genuinely hard reasoning that justified choosing a frontier model in the first place turns out to be a minority of the volume, sometimes a small one.

And every one of those mundane requests is being charged at the rate of the hardest thing your system does, because a default was set during a pilot and defaults are permanent.

This is the most fixable line on most AI bills and it is fixable in weeks rather than quarters. It is also the least glamorous thing in this entire series, which is precisely why it survives for eighteen months in most organizations.

Three routing patterns

Static tiering. The task type determines the model. Classification goes here, extraction goes there, open-ended reasoning goes to the expensive path. Simple, predictable, easy to reason about, easy to explain to a finance function, and it captures the large majority of the available saving. If you do nothing else from this post, do this.

Cascade. Try the cheap tier first. Escalate to the expensive one if some confidence or validation signal says the cheap answer was not good enough. Higher ceiling than static tiering and considerably more complexity, and it lives or dies on one number that I will come to.

Semantic routing. A classifier looks at the request and decides where it goes. Powerful, elegant, and the easiest of the three to over-engineer into something that costs more to operate than it saves. I would reach for this third rather than first, and only once static tiering has told you where the volume actually is.

The cascade math people get wrong

A cascade only saves money if the cheap tier resolves a high enough proportion of requests. Below that proportion, you are paying for two calls instead of one on the escalated portion, plus the added latency, and you have made things worse while feeling sophisticated.

The break-even depends on the price ratio between your tiers. The wider the gap between cheap and expensive, the lower the resolution rate you need to make it worthwhile. With a large ratio, a cheap tier that handles half the traffic is comfortably worth having. With a narrow ratio, you may need it to handle the great majority before the arrangement pays for itself at all.

Work out your own threshold from your own price ratio, then measure your actual escalation rate against it, and put both numbers on the same page.

I would guess that a meaningful share of deployed cascades in production today are sitting below their break-even and nobody has checked, because the cascade was built as an optimization and optimizations do not get audited. They get built, celebrated, and then assumed to be working forever.

Sufficiency, not maximization

The mental shift this requires is the hard part, and it is a business shift rather than a technical one.

The question is not which model is best. That question has a clear answer and it is the wrong question. The question is which model is sufficient for this task, at an acceptable failure rate, given what happens when it fails.

Sufficiency is defined by consequence. A misclassified internal document that a person reviews anyway has a low cost of error, so a high failure rate is tolerable and the cheap tier is fine. A misclassified input feeding an automated decision that reaches a customer has a high cost of error, and the expensive path is correct regardless of unit price.

That is a business judgment. Engineering cannot make it, and when engineering is forced to make it by default, it will reasonably choose the safest option every time, which is how you end up where the jungle was.

You cannot route without evaluation

The dependency that stops most organizations doing any of this.

You cannot route by sufficiency unless you can measure sufficiency. Without an evaluation harness that tells you how each tier performs on each task type, routing is guessing with extra infrastructure, and the first time somebody complains about quality the whole thing gets reverted.

This is the evaluation layer from Lesson 3, the one that appears on nobody’s budget. It turns out to be the precondition for the largest cost lever available. Which is a fairly typical shape in this discipline: the unglamorous investment nobody wanted to fund is what unlocks the saving everybody wants.

Why this belongs at the gateway

Routing must live at a shared gateway rather than inside each application, and this is the architectural decision with the longest tail in the whole series.

If each team makes its own model choice inside its own code, then every choice becomes a hard dependency. Changing tier later means touching forty codebases owned by forty teams with forty roadmaps, which means you will not change tier later, which means every routing decision made this year is permanent.

Put it at the gateway and the tier becomes configuration. You can change it centrally, test it, roll it back, and adjust it every quarter as prices and capabilities move, which they will.

The gateway is also where the tagging from Lesson 1 lives, where the allocation from Lesson 7 becomes possible, and where the guardrails in Lesson 19 will need to sit. It is the single highest-leverage piece of architecture in an enterprise AI estate and it is almost always built late, after the pain, by someone doing a migration.

Three ways this goes wrong

Top tier by default. The pilot configuration, still in production, quietly handling string formatting at premium rates for two years. This is the jungle’s situation and I would expect it to be more common than not.

The negative cascade. Escalation rate below break-even, so the clever thing costs more than the naive thing, and nobody has calculated the threshold to find out.

Routing in the application. Every team choosing independently, every choice becoming permanent, and a tier change turning into a migration project that never gets prioritized.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask two numbers. What share of requests go to your most expensive tier, and what share genuinely require it. The gap between those is money, and it is usually available within a quarter.

If you sit in the Crocodile’s chair, get routing behind a shared gateway before you have forty callers. Retrofitting this is the most avoidable expensive project in the discipline and the window to avoid it closes quietly.

If you sit in the Mandrill’s chair, define the quality floor for your use case, expressed as an acceptable failure rate and what happens when it fails. Sufficiency is a business judgment and if you do not make it, somebody will default to the most expensive option on your behalf, forever.

For everyone: if you run a cascade, calculate its break-even and measure your escalation rate against it. Today. It is a ten minute calculation and there is a real chance it tells you something unwelcome.

Jungle Lesson 12

Most of what you send to the cleverest animal in the jungle does not need the cleverest animal in the jungle. Define what sufficient looks like for each task and buy down to it, because paying for judgment you are not using is the most common and most fixable line on the bill.

Next time: the Beaver takes the Crow back to the river he mentioned all the way back in Lesson 1 and shows her what it is actually carrying. It turns out you are billed for the question as well as the answer, that in most enterprise workloads the question is the larger number, and that every line somebody added to a prompt after an incident two years ago is still being paid for on every single request. Lesson 13 is about context.

The Beaver in your organization has probably already noticed the thing in this post. Whether anyone has asked him is a different question, and it is usually the answer to why it has not been fixed.