Scaling AI FinOps | Lesson 11: Frontier or Local

The Peacock arrived for the second planning cycle with a display of quite exceptional magnificence.

There were slides. There was a total cost of ownership model with a five year horizon and a curve on it that started high and ended gratifyingly low. There was a comparison against what the jungle was currently paying, and the comparison was not close. Somebody in the room said the words game changer without any apparent irony.

The Fox, who had learned a great deal over the preceding two acts, asked a reasonable question about the assumptions.

The Crocodile, who had sat through this presentation four times in his career under four different names and two different technology paradigms, opened one eye and asked the only question that actually decides it.

“What utilization is that curve assuming?”

There was a pause of a length that told everyone in the room what the answer was going to be before the Peacock said it.

Peacocks have enormous displays and modest payloads. This is not dishonesty, it is what the animal is for. The tail is the product. Somebody in the room has to be the one who asks about the bird underneath it, and in most organizations there is no such animal, which is why these decisions go wrong.

Two shapes, not two products

The frontier versus local argument gets conducted as though it were a choice between two products with different features. It is not. It is a choice between two cost shapes, and almost everything else people argue about is downstream of that.

API-served inference is pure variable cost. Near zero fixed. You pay for what you use, forever, and the line is straight. It never gets cheaper through your own effort, only through the provider’s price changes, and it never costs you anything when you are not using it.

Self-hosted inference is high fixed cost with low marginal cost. You pay a large amount whether or not anybody uses it, and then very little per unit on top. The unit economics are terrible at low volume and excellent at high volume, and the entire question is where those two lines cross.

That is the whole thing. Everything else is a variation on it. And it means the honest form of the question is not which is better, it is what volume gets you past the crossover and how confident you are that you will be there.

The utilization trap

Here is where most self-hosted business cases fall over, and it is always the same place.

Fixed-cost economics are quoted at high utilization, because that is the only way the numbers look good, and that is a perfectly reasonable way to describe theoretical capability. They are then delivered at whatever utilization your actual workload produces, which is a completely different number.

Enterprise inference demand is not smooth. It follows the working day, which means it is close to nothing for a substantial portion of every twenty four hours. It follows the working week. It follows the business cycle, the month end, the campaign calendar. It has peaks that determine how much capacity you must provision, and troughs that determine how much of that capacity earns nothing.

You provision for the peak. You are billed for the peak. You utilize somewhere considerably below the average, once you account for the redundancy you need in order to survive the peak going wrong.

An honest utilization assumption for a first-generation enterprise deployment is dramatically lower than the one in the presentation, and the gap between those two numbers is usually the entire business case.

What is missing from the fixed side

The second failure is that the fixed side of the model is systematically incomplete. Four things that belong there and usually are not.

  • The people. Self-hosting is an operational commitment, not a purchase. Somebody keeps it running, patches it, handles the incidents, manages capacity. That is a standing team, and in most models it appears as a footnote or not at all.
  • Redundancy. A single deployment is a single point of failure for something you have just made business critical. The resilient version costs meaningfully more than the demonstration version.
  • Evaluation and update cycles. Every model change you make yourself requires you to re-verify behavior across everything that depends on it. On the API side somebody else absorbs a version of this cost. On your side it is yours, and it recurs.
  • Falling behind. This one has no line item and it is real. Capability moves. A fixed deployment holds still. The gap between what you have and what is available compounds, and at some point it becomes a cost in the form of work your organization cannot do.

The three questions

What I would ask, in this order, of any proposal in either direction.

What is the break-even volume? A specific number of units per period at which the two curves cross, with the fully loaded fixed side included. If nobody can produce this number, the analysis has not been done, whatever else has been produced.

How confident are you in the volume forecast? Not what is the forecast. How confident. Because the forecast is doing all the work and it is usually the least examined input in the model. Ask when it was made, by whom, and what it is based on.

What does this look like at half that volume? This is the question that decides it. Almost every self-hosting case survives the first two questions and fails the third, because the fixed side does not move and the unit economics collapse. If the answer at half volume is still acceptable, you have a genuinely robust case. If it is catastrophic, you are making a bet on a forecast rather than an architecture decision.

What optionality is worth

The thing that is almost never priced, and which I think matters more than anything else on this list.

Variable cost buys you the right to change your mind. Fixed cost buys you a lower unit price and a commitment to a capability level that will be surpassed, on a timeline you do not control, while you are still depreciating the thing.

In a slow-moving field that trade is usually worth taking, which is why organizations have made it comfortably for decades with every other kind of infrastructure. In a field where capability and price are both moving quickly, the option to move has genuine value, and the discount you are being offered has to exceed it.

The framing I would use in the room: what would you pay, today, for the ability to shift this entire workload somewhere else in six months with no stranded cost? Whatever that number is, it is the hurdle the discount has to clear. Nobody ever calculates it, and it is frequently larger than the saving.

The answer is usually both

For most enterprises the correct answer is not one or the other, and treating it as a single organizational decision is how it goes wrong.

Work that is high volume, narrow, well specified, stable in its requirements, latency sensitive, or subject to a constraint on where it can run, suits a fixed-cost local deployment. The volume is there, the capability requirement is not moving, and the crossover is genuinely favorable.

Work that is low volume, broad, hard, exploratory, or benefits from the newest capability suits variable-cost API-served inference. You are buying capability and optionality rather than unit price, and both are worth the premium.

The organizations that get this right decide per workload and build the routing layer that makes the decision reversible. That routing layer is the next lesson, and it is the mechanism that turns this from a one-time architectural bet into something you can adjust quarterly.

Three ways this goes wrong

The full utilization business case. Economics quoted at a utilization the workload will never reach, delivered at half of it, defended for two years by people who cannot admit the input was wrong.

The forgotten team. Hardware and capacity modeled carefully, the standing operational function required to run it modeled not at all. This single omission has sunk more self-hosting cases than every other factor combined.

Ideological procurement. The architecture chosen first, for reasons of principle or preference, and the model constructed backwards to support it. This happens in both directions and it is equally expensive either way.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask the three questions in order and do not accept a break-even volume without a stated confidence in the forecast underneath it. Then ask for the version at half volume before you approve anything.

If you sit in the Crocodile’s chair, model the fully loaded fixed side including people, redundancy, evaluation cycles and refresh. You are the only animal in the building who can produce that number honestly, and it is the number that decides it.

If you sit in the Mandrill’s chair, be honest with yourself about whether your volume forecast is a forecast or an aspiration. The architecture decision is far more expensive to reverse than the forecast is to correct.

For everyone: decide per workload, never once for the organization. A single global answer to this question is always wrong somewhere, and it is usually wrong somewhere expensive.

Jungle Lesson 11

This is not a choice between two products, it is a choice between two cost curves, and the only questions are where they cross and how much you trust the volume that gets you there. Variable cost is the price you pay for the right to change your mind, and in a field moving this fast that right is worth more than most business cases give it credit for.

Next time: the Beaver builds something over a weekend that takes a large bite out of the bill, and nobody notices for a month because the outputs did not change at all. The Fox is mildly annoyed that the biggest win of the quarter came from plumbing rather than strategy. Lesson 12 is about the routing layer, and why the cheapest sufficient answer beats the smartest one.

If a proposal in either direction cannot tell you what it looks like at half the assumed volume, it is not an architecture decision. It is a forecast with an architecture attached, and forecasts are much cheaper to be wrong about.

Scaling AI FinOps | Lesson 10: The Value Ledger

The Tortoise had been reading the Fox’s revised benefits case for some time before she said anything, which is normal for a tortoise and unnerving for everyone else.

“You know what this is,” she said eventually.

The Fox said that it was a benefits case.

“It is an audit trail,” said the Tortoise. “You have written down what you claimed, how you intend to prove it, who is accountable, and when it lands. That is the document I have been asking this jungle for since before any of you had heard the word inference. You have built it by accident while trying to survive a finance meeting.”

The Fox, who had spent a considerable portion of the previous two seasons regarding the Tortoise as an obstacle, took a moment with that.

“Are we on the same side?” he asked.

“We have been the entire time,” said the Tortoise. “You were busy.”

This is the alliance that ends up mattering most in the rest of this story, and I want to flag it now because it takes most organizations far too long to discover. The compliance function and the AI function want the same artifact for different reasons. One of them wants to be able to defend a decision to a regulator. The other wants to be able to defend a decision to a CFO. It is the same document.

What a value ledger actually is

A standing record, one row per claimed benefit, carrying at minimum:

  • What was claimed, in a unit somebody can check
  • Which conversion path it uses, from Lesson 8
  • Who owns realizing it, by name rather than by function
  • How strong the evidence is, from Lesson 9’s ladder
  • When it is expected to land
  • What actually landed, filled in afterwards, next to the claim rather than replacing it

That is the whole instrument. It is not sophisticated. It could live in a spreadsheet and in most organizations it should, at least for the first year, because the discipline is the hard part and the tooling is not.

Append only, and why that is the entire point

One rule matters more than all the others combined. Entries are never edited. What was claimed stays visible next to what happened.

The temptation to restate is enormous and it always arrives dressed as accuracy. We understand the process better now. The original definition was flawed. The scope changed. Every one of these is often true, and every one of them, acted upon, destroys the only thing the ledger was for.

Because the value of the instrument is not the record of benefits. It is the record of the gap between what people predicted and what occurred, accumulated over enough cycles that the organization learns the shape of its own optimism. That is genuinely valuable. A team that has systematically overclaimed by a factor of three for six quarters is a team whose next claim you can adjust with confidence, and that adjustment is worth more than any estimation methodology you could buy.

Edit the entries and you have a document that shows every claim was roughly correct, which teaches nothing and predicts nothing. If something needs restating, add a new row with a new date and leave the old one where it is.

Benefits decay

The second discipline, and the one that catches organizations in year two, which is where Act V of this series is heading.

Most AI benefits are not annuities. They erode. The process changes and the saving no longer applies to a process that exists. Volumes shift. The comparison process itself improves for unrelated reasons, so the delta narrows. Staff turn over and the new people never did it the old way, so the improvement is invisible to them and stops being defended.

Booking a benefit once and carrying it forward indefinitely is the single most common overstatement in this discipline, and it is almost never deliberate. It happens because there is no field in the document that says when this stops being true.

So add one. Every entry carries an explicit decay assumption and a review date. Flat for three years, or declining after eighteen months, or one-off. You will be wrong about the assumption. Being wrong about a documented assumption is a vastly better position than being silently wrong about an undocumented one, because the first one gets corrected on a date and the second one gets discovered in an audit.

Evidence grading

Every entry carries a grade, taken directly from Lesson 9’s ladder. Holdout, staggered rollout, adjusted pre and post, cohort comparison, self-report.

The purpose is not to disqualify weak evidence. Weak evidence is normal and often it is all you can get. The purpose is proportionality. A claim supported by a survey is still a claim, it simply does not get to fund a decision that requires a claim supported by a holdout.

This does something subtle and useful to organizational behavior. Once the grade is visible next to the claim, teams start volunteering to improve their evidence, because a stronger grade means their number carries more weight in the reallocation conversation. You have made rigor competitively advantageous rather than a compliance burden, which is the only way it ever actually happens.

The reconciliation habit

Once a cycle, quarterly for most organizations, put claimed next to realized and publish the variance.

The first time you do this it will be uncomfortable. The second time it will be uncomfortable. Somewhere around the third or fourth cycle something changes, and the document becomes the most trusted artifact the program produces, because it is the only one that has ever voluntarily reported its own misses.

I have watched this transformation happen and it is worth the two bad quarters. A program that publishes its own variance gets believed about everything else. A program that only ever reports successes gets discounted on everything, including the successes, and the discount is applied by people who will never tell you they are applying it.

The Crow’s observation, when the first reconciliation came in and the realized number was about a third of the claimed one: “This is the first document anyone has given me about this subject that I would repeat to the board without checking it first.”

This is governance, not reporting

The ledger is what turns AI FinOps from a reporting function into a governance one, and this is the thing I would most want a leadership team to understand from Act II.

Reporting tells you what happened. Governance changes what happens next. A ledger with owners, dates, decay assumptions and published variance does the second thing, because it creates a moment where somebody has to account for a prediction they made, in front of people who can see the original.

It is also, as the Tortoise spotted immediately, exactly what external scrutiny asks for first. Not your architecture. Not your model choices. What did you claim, how did you verify it, who was accountable, and what happened. Every serious review of an AI program I have seen has converged on that set of questions, and the organizations that had a ledger answered them in an afternoon.

And it is what makes the chargeback conversation from Lesson 7 survivable, because a charge without a ledger is a bill, and a charge with a ledger is a price for something whose value has been demonstrated.

Three ways this goes wrong

The retroactive edit. Last year’s claim quietly aligned to this year’s outcome, for excellent reasons, destroying the instrument entirely while making it look tidier.

Perpetual benefits. Booked once, assumed forever, never reviewed, still sitting in the cumulative total four years later describing a process that was replaced twice.

The ledger nobody reads. Maintained diligently by someone conscientious, reviewed by no body with authority to act, quietly reclassified as a compliance artifact. This is the most common failure and it is entirely a governance design problem rather than a documentation one.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, own the ledger and make append-only a written policy rather than an aspiration. This is a finance instrument. If it lives in the technology function it will be maintained beautifully and read by nobody.

If you sit in the Crocodile’s chair, feed it automatically wherever the data allows. Manually maintained ledgers decay within two quarters regardless of how committed everybody was at the start, and the decay is always discovered at the worst moment.

If you sit in the Mandrill’s chair, put your name against your realization dates. The accountability is the point and it is also the fastest route to being trusted with a larger budget, because you will be one of very few people who can point at a prediction they made and hit.

For everyone: publish the claimed versus realized variance. It is the single most credibility-generating document available to an AI program and almost nobody produces it.

Jungle Lesson 10

Write down what you promised, who owns it, and when it lands, then never edit the entry. A benefits case that cannot be checked against what actually happened is not a case, it is a mood, and moods do not survive the second budget cycle.

That closes the second act. The jungle can now count. It has a denominator it chose rather than inherited, an allocation rule it published before it published a number, an honest view of what productivity gains actually convert into money, a way of proving things that does not rely on memory, and a ledger that records what was promised next to what arrived.

None of which has yet made anything cheaper.

That is the third act, and it moves the argument somewhere most finance functions do not expect it to go, which is into the architecture. Almost every meaningful lever on AI unit economics is an engineering decision made months before anyone looks at a bill. Where inference runs, which model handles which task, how much context gets sent, where the data sits, and how much autonomy the system has. Act III opens with the one everybody wants to argue about, which is whether to run it yourself. Lesson 11 does the crossover math honestly.

If your organization has a compliance officer who has been asking for something like this for two years, the fastest route to a value ledger is to go and find them. They have thought about it more carefully than you have and they have been waiting for somebody to want it.

Scaling AI FinOps | Lesson 9: Baselines, Counterfactuals and Holdouts

After the meeting in the last lesson, the jungle needed evidence, and discovered it did not have any.

Not because nobody had measured anything. Everybody had measured a great deal. The problem was that all of it had been measured after the capability went live, which meant every comparison was against a version of the old process that existed only in the memories of the animals who had stopped doing it.

It was at this point that the Owl said something that changed the quarter.

“I still have the north territory,” he said.

Everyone looked at him.

What had happened, it emerged, was this. When the rollout was designed, the Owl had produced a phased plan, in the way that Owls do, with the north territory scheduled last for reasons involving a system migration. The migration had been delayed. Nobody had revisited the plan, because nobody revisits a rollout plan once the rollout is going well. So the north territory had spent two full seasons doing the old process, unchanged, while everybody else changed, and their operational data had been collecting the whole time in the ordinary way.

He had, entirely by accident and with no intent whatsoever, been running a control group for eight months.

“Did you know?” asked the Crow.

“I knew they hadn’t gone live,” said the Owl carefully. “I did not know that was interesting.”

Owls have very large eyes. On this one occasion, something had actually been in front of them.

The remembered baseline

Almost every organization begins measuring after deployment. This is not laziness. It is that before deployment nobody yet cares, and by the time somebody cares, the thing they wanted to compare against has been replaced.

So the comparison becomes a comparison against memory, and memory has a consistent bias in this specific situation. People remember the old process as slower and more painful than it was, because what they remember most vividly are the bad instances. Ask a team how long something used to take and you will reliably get a number closer to the worst case than the median.

Which produces benefit cases that are inflated in a direction nobody chose, built by people acting in good faith, and completely undefendable the moment anyone asks where the baseline came from.

The baseline is the whole game. Everything else in this post is technique for when you did not capture one, which is most of the time.

The methods ladder

Four approaches, strongest first. Use the highest rung you can reach.

Randomized holdout. A comparable population that does not get access, chosen at random, for a defined period. This is the cleanest thing available and it is politically the hardest, which is why it is rare. It is also possible for far longer than most organizations assume, because full deployment usually takes longer than anyone plans anyway.

Staggered rollout. Territories go live at different times, and the not-yet-live ones serve as the comparison. This is nearly free, because you are almost certainly already doing it, and the great majority of organizations throw the resulting evidence away by not capturing it. The Owl’s accidental control group was this, discovered eight months late.

Pre and post with controls. Before and after, adjusted for the confounders you can identify. Seasonality, volume changes, staffing changes, process changes shipped at the same time. Weaker than the first two but genuinely usable if you are honest about the adjustments.

Cohort comparison. Heavy users against light users. This is the weakest rung and the most commonly used, because it requires nothing but the data you already have. The problem is that self-selection is doing most of the work. The animals who adopted heavily are not a random sample. They are the ones who were already better at this, or more motivated, or had the kind of work that suited it, and you are measuring that rather than the capability.

The confounders that bite

Four that show up specifically in this domain and that will eat an otherwise sound analysis.

  • The process change that shipped alongside. Almost nobody deploys a capability without also simplifying the workflow around it. That simplification often produces more of the benefit than the capability does, and the two are attributed together.
  • Novelty. New things get attention and effort. Some of the early improvement is people trying harder because something is new, and it goes away.
  • Observation. Teams that know they are being measured perform differently. This is old news and it applies here as much as anywhere.
  • Adoption order. Your best people adopt first, reliably. Early results are therefore measured on your strongest performers and will not replicate across the rest.

Making a holdout survivable

The objection to holdouts is always the same and it is not unreasonable. You are denying a useful tool to a group of colleagues in order to generate a chart.

The framing that works, and I have seen it work several times, is to stop describing it as denial. The holdout group is not being deprived. They are the reason the program can prove it works, which is the reason it will still be funded next year, which is the reason everyone eventually gets it. That is true and people respond to it being said out loud.

Three things make it stick. Timebox it, with a date. Guarantee access at the end, in writing. And name the group publicly as the reason the evidence exists, in the same presentation where the results appear. The Mandrill volunteered one of his teams for this in the following quarter, and it did him considerably more good politically than it cost him operationally.

When to measure

Three points, and you need all three.

Week two, which is your peak and which you should record while knowing it is a peak. Month three, which is where the novelty has gone and the real shape starts to appear. Month twelve, which is the number that should drive any permanent funding decision.

The numbers will fall between these points and that is not failure, it is the reference point moving as the improved process becomes normal. What matters is that the organization sees all three, because an organization that only ever sees week two will make every decision on a peak, and an organization that only sees month twelve will conclude the whole thing underdelivered without ever knowing what the early state looked like.

Put the month twelve measurement in the calendar now, while everybody is still enthusiastic enough to agree to it. Nobody schedules that meeting later.

How much rigor is enough

A brief word against overdoing this, since I have just spent a thousand words arguing for evidence.

Match the strength of your evidence to the stakes of the decision. A modest capability serving one team does not need a randomized holdout. Pre and post with a sensible adjustment is fine, and demanding more is a way of making sure nothing gets measured at all because the bar is exhausting.

A decision about whether to deploy something across the organization permanently is a different matter. There the cost of a holdout is trivial against the cost of being wrong, and the fact that it is inconvenient is not an argument.

The failure mode I see most is uniform rigor. Either everything gets the full treatment, which means nothing does, or nothing does, which means everything is anecdote. Tier it.

Three ways this goes wrong

The remembered baseline. Comparing against how long people think it used to take. Ubiquitous, invisible, and fatal the moment anyone asks.

Rollout without capture. Running a staggered deployment, generating a natural experiment for free, and recording none of it. This is the most wasteful thing in this post because the evidence was already being produced.

Measuring once, at the peak. Funding a permanent capability on a week two number, then spending year two explaining why the benefits shrank when nothing shrank except the novelty.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask when the baseline was captured and by whom. If the answer is after go-live, discount the case accordingly and say so out loud, so that the next person to present to you captures one first.

If you sit in the Crocodile’s chair, capture the staggered rollout data. You are already generating it. Recording it costs close to nothing and it is worth more than any dashboard you could build this year.

If you sit in the Mandrill’s chair, volunteer a team as a holdout and get named publicly as the reason the program can prove anything. It is a much better position than it appears from the outside.

For everyone: put the month twelve measurement in the calendar today. It is the only one of the three that nobody will schedule voluntarily and the only one that should drive a permanent decision.

Jungle Lesson 9

You cannot measure a change you did not measure before. Capture the baseline while you are still bored by the idea, because by the time you need it, everyone will remember the old process as being much worse than it was, and you will have no way to prove otherwise.

Next time: the Tortoise looks at everything the Fox has built over the last four lessons and points out, with some satisfaction, that he has accidentally constructed the audit trail she has been asking for since before any of this started. Compliance and the AI lead end up on the same side of the table, which surprises both of them. Lesson 10 closes the second act with the value ledger.

Somewhere in your organization there is probably an Owl sitting on a delayed rollout that nobody has revisited. That is not a project management failure this month. It is the only clean comparison you are going to get.

Scaling AI FinOps | Lesson 8: The Hours Saved Fallacy

The number on the Fox’s slide was very large, and he had checked it three times.

Six hundred and eighty animals were using the capability. Survey responses indicated an average of twenty two minutes saved per animal per day. Multiply that out across a working year, apply a loaded hourly rate, and you arrive at a figure with a great many digits in it, sitting in a box, in bold, next to a modest and considerably smaller number representing what the whole thing cost.

It was, by the standards of these presentations, a good one. The arithmetic was correct. The survey was real. The adoption numbers were not inflated. I have seen a hundred versions of this slide and this was among the more honest ones.

The Crow looked at it for a while.

She did not challenge the twenty two minutes. She did not question the adoption figure or ask about the survey methodology, which is what the Fox had prepared for and which would have been a conversation he could win.

“Where did the time go?” she asked.

The Fox said that it went back to the animals.

“Yes,” said the Crow. “And then what?”

The room was quiet for long enough that everybody in it understood what had just happened, including the Fox, who to his considerable credit did not attempt to fill the silence with anything.

The Hyena, for once, was not laughing. “That’s the whole thing, isn’t it,” she said. “You’ve proved they have more time. You haven’t proved you have anything.”

Why twenty minutes is not twenty minutes

Here is the fallacy, stated plainly, because it is the single most successful piece of fiction in modern enterprise reporting and it is almost never named directly.

Twenty two minutes saved by each of six hundred and eighty animals is not roughly two hundred and fifty person days of recovered capacity. It is six hundred and eighty animals with slightly better afternoons.

Those are not the same thing and they are not close to the same thing. The first is a resource the organization can deploy. The second is a genuinely nice outcome that produces no line in any ledger, ever, under any circumstances.

Time saved becomes money only when it is aggregated into something the organization can either redeploy or avoid paying for. Twenty two minutes, distributed across six hundred and eighty people, in fragments, throughout a day, aggregates into nothing at all. It is real. Every one of those animals genuinely got the time back. It simply does not add up into a unit anybody can spend.

The three conversion paths

There are exactly three ways time turns into money, and I would rank them by how well they survive scrutiny.

Cost avoidance. A role you did not need to backfill. A contractor whose renewal you did not sign. Overtime you did not pay. Seasonal capacity you did not bring in. This is the strongest by a wide margin, because it is verifiable in the ledger by someone who was not involved in the project. If you can point at a requisition that was cancelled, you have a benefit, and nobody can argue with you.

Capacity redeployment. The same people producing measurably more of something the organization values. Not more availability. More output. This is real and it is the most common genuine benefit, but it requires an output metric rather than an input one. “The team has more time for strategic work” is not this. “The team closed forty percent more cases” is.

Quality and cycle time. Faster resolution, fewer errors, better retention, higher conversion. Frequently the largest of the three in absolute terms, and the hardest to trace, because it needs a chain of evidence running from the capability all the way to something financial. Worth doing. Not worth pretending is easy.

And then the part people find uncomfortable. Anything that does not land in one of those three is not a benefit. It is a nice thing.

Nice things matter. Less frustrating work, less time on tasks people hate, better morale, lower attrition risk. I am not dismissing any of it and some of it is genuinely valuable. But it should be reported as what it is, in its own section, under its own heading, rather than converted into currency by multiplication and placed next to the cost line as though the two numbers were the same kind of object.

The fragmentation threshold

Underneath all of this sits a threshold effect that I think deserves more attention than it gets.

Below some proportion of a role’s time, savings do not aggregate at all. They dissipate. The person absorbs the time into the general texture of their day and nothing further happens, because there is no mechanism by which fifteen scattered minutes become a deployable resource.

Above some proportion, savings become visible enough that the work can be restructured, and restructuring is what actually converts time into capacity or cost. A role that gets thirty percent of its time back can be redesigned. A role that gets four percent back cannot.

Where exactly that threshold sits depends on the work, and I would be suspicious of anyone quoting you a universal number. But the existence of the threshold is the important part, because it explains something that otherwise looks like a paradox: why broad shallow deployments produce enormous claimed benefits and almost no realized ones, while narrow deep deployments produce modest claimed benefits that actually show up.

If your deployment strategy is a small saving for a very large number of people, you have chosen the shape that maximizes the claimable number and minimizes the realizable one. That may still be the right choice. It should be a conscious one.

Self-reported savings, and what happens to them

Two things about survey data, since it is the evidence almost everyone uses.

First, people overestimate time saved, consistently and without any intent to mislead. Asking someone how much time a tool saves them is asking them to compare their current experience against a remembered version of a process they have stopped doing. Memory is generous about this in a fairly reliable direction.

Second, and more usefully, reported savings decay. Measure at week two and you get a peak, because the contrast with the old way is vivid and the novelty is real. Measure the same population at month six and the number will be materially lower. Measure at month twelve and lower again, as the improved process becomes the baseline against which nothing feels saved.

This is not people becoming disillusioned. It is the reference point moving, which is exactly what you wanted to happen. But it means an organization that measures once, at week two, and then funds a permanent capability on that number, has funded on a peak that will never recur.

The honest format

What I would put on the slide instead, and what the Fox put on the next version of his.

State the gross number. Twenty two minutes, six hundred and eighty animals, here is what that multiplies to. Do not hide it, it is real and it is the reason anyone is interested.

Then state the conversion path for each portion of it. This much lands as cost avoidance, here is the specific requisition. This much lands as redeployment, here is the output metric that will show it. This much has no conversion path and is reported as a nice thing.

Then state the realized number, which will be dramatically smaller. And then show the gap between the two, explicitly, as its own line.

That gap is not a weakness in your case. Showing it is the single thing that makes everything else in the presentation believable, because it demonstrates that you understand the difference between the two numbers, which is the exact thing the Crow was testing for when she asked her question.

The Crow is not the enemy

One more thing, aimed at anyone who reads this and feels defensive on the Fox’s behalf.

A benefits case that survives scrutiny gets funded again next year. A benefits case that does not survive scrutiny does not merely fail, it takes the credibility of the whole program with it, and the next request from the same team starts from a worse position than the first one did.

The Crow was not trying to kill the capability. She was trying to find out whether she could defend it to somebody more senior than her, six months from now, in a room the Fox will not be in. The question she asked is the question she will be asked. She was doing him a favor and it took him about a week to work that out.

Three ways this goes wrong

Loaded rate multiplication. Hours times salary, no conversion path, presented with total confidence. This works exactly once, on an audience that has not seen it before.

Double counting. The same saved hours claimed by the platform team, the use case owner, and the enterprise transformation program, none of whom are aware the others are claiming them. In a large organization this is startlingly common and it only surfaces when somebody adds up all the claimed benefits and finds they exceed the total cost base of the function.

The benefit with no owner. Claimed in the case, never assigned to anybody, therefore never realized, and nobody notices because nobody was watching for it. This is the most common of the three and the quietest.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask two questions of every claimed benefit. Which of the three conversion paths does this use, and who owns realizing it. Thirty seconds, and it will do more for the quality of your portfolio than any amount of process.

If you sit in the Crocodile’s chair, give the business the usage and acceptance data they need to make an honest case. Withholding it does not protect the program. It guarantees the case gets built on survey data instead, which is weaker, and which will fail later at higher cost.

If you sit in the Mandrill’s chair, commit to a realized number rather than a gross one, and accept that it will be a fraction of the headline. A smaller number you hit is worth considerably more to your standing than a large one you miss, and it is the only version that gets you funded a third time.

For everyone: stop reporting hours saved as a headline figure. Report the conversion. If there is no conversion, report it as a nice thing, in its own box, honestly labeled.

Jungle Lesson 8

Time saved is not money saved until somebody does something specific with the time. Twenty minutes back for six hundred animals is not two hundred days of capacity, it is six hundred slightly better afternoons, and you cannot put an afternoon in the ledger.

Next time: the Owl finally has his moment. It turns out he has been quietly holding a control group for two seasons because nobody ever told him to stop, which makes him the only animal in the jungle who can prove anything at all. Lesson 9 is about baselines, counterfactuals and holdouts, and how to demonstrate value when you cannot run a clean experiment.

If you are about to present a benefits case built on survey data and a loaded hourly rate, the question at the top of this piece is the one you will be asked. It is worth having an answer before somebody else has the silence.

Scaling AI FinOps | Lesson 7: Who Pays for the Watering Hole

The argument about who paid for the Watering Hole took four meetings, and it was the first time in this whole story that nobody was confused about anything.

That is what made it different. Every previous fight had been a misunderstanding at heart. This one was not. Every animal in the room understood the question perfectly, understood what the other animals wanted, and understood exactly why they wanted it. They just disagreed, in the ordinary way, about who should carry a cost that all of them benefited from.

The Mandrill made the strongest case. He pointed out that his territory had been the first to adopt, that early adoption had subsidized the platform’s development for everyone else, and that charging him now on consumption would penalize him for having taken the risk while the cautious territories waited.

This was completely true. It was also, and he knew it, entirely beside the point, because the same argument would be equally available to him next year and the year after. The Hyena, from her branch, described it as the finest piece of reasoning she had ever heard deployed in service of not paying for something.

The Sloth arrived at the fourth meeting with a framework. It was a good framework. It would have been extremely useful at the first meeting, roughly six weeks earlier, which the Sloth acknowledged without any evident distress.

“The destination,” he said, “was always correct.”

This is not a technical problem

I want to be blunt about this because I have watched a lot of organizations get it wrong in the same way. Allocation is not a tooling problem, it is not a data problem, and it will not be solved by buying something.

It is a question about how your organization distributes the cost of a shared good among parties who all want it to exist and all prefer someone else to fund it. That is a political economy question. It has been a political economy question for as long as organizations have had shared services, and the fact that this particular shared service involves inference does not change its nature at all.

Which means the answer has to be decided, announced, and defended. It cannot be discovered in a dashboard.

Three models, and what they actually do

Central absorption. The platform sits on a central budget and is free at the point of use. Adoption is fast, because nothing is easier to adopt than something free. Consumption discipline is precisely zero, because there is no reason for it to exist. And by the second year the central line has grown large enough that somebody senior asks what it is, at which point it is politically indefensible, because nobody can attribute a single unit of it to a single business outcome.

Showback. Costs are attributed and reported to consuming territories but not actually charged. Behavior changes, genuinely, but modestly. People do respond to seeing their own number even when nothing happens as a result of it. This is the right starting position for almost every organization and it is treated as a waypoint when it should be treated as a destination for at least a year.

Chargeback. Costs move to the consuming budget. Discipline becomes real, immediately and unmistakably. So do the perverse incentives, which arrive in the same week and which almost nobody plans for.

What chargeback does to behavior

Chargeback is not wrong. It is the correct end state for a mature platform. But it changes what people optimize for, and the changes are worth naming before you switch it on rather than discovering them afterwards.

  • Local optimization beats global. A territory will happily take an action that reduces its own charge and increases total organizational cost, because the first number is on its report card and the second one is not.
  • Avoidance of the sanctioned path. If using the platform incurs a charge and using something else does not, some proportion of your organization will use something else. You have just recreated the conditions that produced the Tortoise’s list two lessons ago.
  • Under-investment in the commons. Nobody volunteers to fund an improvement to a shared component that mostly benefits other territories, so shared components stop getting improved.
  • Experimentation stops. This is the one that costs the most and shows up the least. Exploration is cheap in absolute terms and highly visible on a chargeback line, so it is the first thing a territory cuts when it wants its number down.

The shared component problem

Underneath the political argument sits a real structural one that pure consumption splitting cannot handle.

Some things in the Watering Hole serve everybody and are attributable to nobody. The retrieval index. The evaluation harness. The guardrails. The gateway. These exist so that any capability can be built at all, and their cost has almost nothing to do with how much any particular territory consumes.

Split those on consumption and you get an absurd outcome, which is that the first territory to adopt pays for digging the hole and every subsequent one drinks from it at marginal cost. This is not a hypothetical. It is exactly what the Mandrill was complaining about, which is why his argument was annoying rather than wrong.

The Crocodile, who has watched this pattern play out with every shared platform of the last twenty years, offered the observation that ended the fourth meeting. “You are trying to solve two problems with one rule,” he said. “Stop.”

The shape that works

Two mechanisms, not one.

A fixed floor, funded centrally, covering the shared foundation. The gateway, the retrieval infrastructure, the evaluation capability, the guardrails, the people who keep it running. This is infrastructure. You do not consumption-split the electrical system.

A variable band above it, tracking actual consumption, shown back initially and charged back once the numbers are trusted. This is the part that should respond to behavior, because it is the part behavior actually drives.

And a published rule defining which is which. Here the rule matters considerably more than where the line falls. Almost any defensible boundary will work if it is written down, explained, and applied consistently. No boundary will work if it is discovered by territories one invoice at a time.

Tagging is the precondition

None of this functions on untagged traffic, which is why the Field Kit item back in Lesson 1 was request level instrumentation and why I have kept returning to it.

If you cannot say which team, which use case and which business purpose generated a given unit of consumption, you cannot allocate. You can estimate, and estimates get disputed, and disputed allocations produce meetings about attribution methodology instead of meetings about decisions.

The unglamorous truth is that the organization that did three weeks of tagging work early can have this entire conversation, and the one that did not cannot, regardless of how sophisticated its thinking about allocation models happens to be.

The maturity path

Absorb, then show back, then charge back. In that order, and do not skip.

The skip is tempting because chargeback is obviously the mature answer and absorption is obviously the immature one, so why spend a year in the middle. The reason is trust. Chargeback introduced before anyone believes the numbers converts every review into a dispute about attribution accuracy. You will spend a year arguing about the data instead of a year making decisions with it, and at the end of that year the credibility of the whole exercise will be lower than when you started.

Showback is where the numbers get argued into shape at low stakes. That is not a delay. That is the work.

Three ways this goes wrong

Chargeback before trust. Real money moving on numbers nobody believes. Every meeting becomes a methodology seminar.

The free lunch that ends abruptly. Central absorption in year one, sudden chargeback in year two, adoption falls off a cliff, and the drop gets reported upward as evidence that the business was not really interested after all.

Precision theater. Enormous effort spent allocating the last few percent accurately while the entire shared foundation sits in an unexamined central bucket that nobody has looked at in three quarters.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, publish the allocation rule before you publish the first allocated number. People can accept or reject a rule. They cannot do either retroactively, and if the rule arrives after the invoice it will be read as a justification rather than a policy.

If you sit in the Crocodile’s chair, make the tag mandatory at the gateway. Untagged traffic gets rejected, not absorbed. This will be genuinely unpopular for about two weeks and it settles the question for about two years.

If you sit in the Mandrill’s chair, have the argument about the model now, loudly, while it is still abstract. Once the numbers exist, every argument you make about the model will be heard as an argument about your own bill, and you will lose it on those grounds regardless of whether you are right.

For everyone: name the shared components explicitly, in a list, and decide who funds them separately from the consumption question. Two problems, two mechanisms.

Jungle Lesson 7

Everybody wants the watering hole and nobody wants it on their books. Decide who pays for the shared foundation before you decide how to split the water, because a consumption split with no floor makes the first animal to drink pay for digging the hole.

Next time: the Fox presents a benefits case with a very large number in it. The Crow does not challenge the number. She asks one question about where the hours actually went, and the room goes quiet for long enough that everyone understands what has just happened. Lesson 8 is about quantifying productivity without lying, and it is the one I would most like people to read.

The Sloth’s framework, incidentally, was good. It arrived late because nobody asked him until the situation was already unpleasant. There is usually a Sloth, and there is usually a framework, and the delay is more often a demand problem than a supply one.

Scaling AI FinOps | Lesson 6: Choosing a Denominator

The rains came, which in the jungle means budget season, and the Crow did something the Fox had not been expecting.

She offered him a deal.

“Pick your denominator,” she said. “Whatever unit of work you think this capability produces. You choose it, not me. I will fund against it and I will not argue with the one you pick.”

The Fox, who had spent two seasons being told that finance did not understand what he was building, waited for the catch.

“You own it for two years,” said the Crow. “You do not get to change it when it stops flattering you.”

That was the catch. It was also, though it took him most of a month to see it, the fairest offer anybody had made him since this began.

He had assumed choosing would take an afternoon. It took three weeks, and he abandoned four candidates before he found one he was willing to sign his name to for two years. Somewhere in the middle of the second week he came to a realization that I think most people in his position eventually arrive at, usually later than he did.

Complaining about not having a denominator is enormously easier than choosing one, and he had been doing the easy version for a year.

Why choosing is hard

The four tests from Lesson 2 still apply. Countable without new instrumentation, meaningful without explanation, attributable to someone, stable across quarters. Those get you a shortlist.

What they do not tell you is the thing that makes the decision genuinely difficult, which is that every honest denominator makes your capability look worse than the dishonest one sitting next to it. Choosing well is an act of deliberately taking a smaller number, on purpose, and defending it. That is why so few organizations do it, and it is not because they lack a framework.

The gap between generated and accepted

Here is the distinction that does most of the work in this post.

Cost per generated output is easy to measure and almost useless. Every system produces output. Producing output is the one thing you can absolutely rely on it to do, whether or not the output was any good, whether or not anyone read it, whether or not it went anywhere.

Cost per accepted output is hard to measure and true. Accepted meaning used. Sent. Merged. Approved. Acted upon. The thing where a person looked at what came out and decided it was good enough to carry forward.

Almost the entire value question lives in the gap between those two numbers, and organizations that measure the easy one systematically overstate their own performance, sometimes by a lot. I have seen acceptance rates that would have changed a funding decision if anyone had thought to look, sitting quietly inside a program reporting excellent volume metrics every month.

The unpleasant part is that acceptance is often the one thing nobody instrumented, because at pilot stage it did not matter. Everything was accepted at pilot stage. That was rather the point of a pilot, and it is why pilot metrics travel so badly into production.

Four archetypes

What this looks like in practice, with the acceptance definition attached, because the definition is the hard half.

  • Service and support work. Cost per resolved case. Accepted means the case closed and did not reopen within some window you choose in advance and then do not adjust.
  • Content and communication. Cost per accepted draft. Accepted means it went out, or went to the next stage, with or without editing. If you want to be rigorous, track the edit distance too, because a draft that gets rewritten entirely was not accepted, it was raw material.
  • Operations and processing. Cost per processed document. Accepted means it cleared without human correction. This is the cleanest of the four and the one most organizations can measure today if they look.
  • Engineering work. Cost per merged change. Accepted means it went into the main line. Suggestions that were dismissed are not free and should not be invisible.

Notice that in every case the acceptance definition contains a judgment call, and the judgment call belongs to the business rather than to engineering. This is why the Mandrill has to be in the room for this conversation, and why it goes badly when he is not.

The two traps

Two denominators are chosen constantly and are wrong in ways worth being explicit about.

Cost per user. This is the most common choice and the most reliably damaging one, because it rewards exactly the wrong behavior. Under cost per user, the way to improve your metric is to add more users. Adding users who barely touch the thing improves your number considerably. Meanwhile your heaviest users, the ones actually generating the value, make the metric look worse, which creates a quiet institutional pressure to discourage the people you most want to encourage. It is the metric that most reliably produces the wrong decision, and it is on more slides than any other.

Cost per token, or per request, or per call. This is a genuinely useful efficiency metric and it belongs to the Crocodile, in his own meetings, where it will do good work. It is not a business metric and it should not be near a steering committee, because it tells you how efficiently you are doing something without telling you whether the something was worth doing. A system can get steadily cheaper per request while producing steadily less value, and cost per request will report that as a success story every single month.

Does it survive a bad quarter

The test I would add to the four from Lesson 2 is this one, and it is the one that catches the metrics chosen for the wrong reasons.

Imagine the capability has a bad quarter. Volume down, quality complaints up, a territory unhappy. Does your denominator show that? Or does it stay flat, or improve, because of some property of how it is constructed?

If a metric only looks sensible when things are going well, it is not a measurement. It is a narrative device, and the moment it is needed it will not be there. The Fox discarded two of his four candidates on precisely this test, which is why the three weeks were well spent.

One more thing worth saying. More than one denominator per capability is fine, and often better, because different consumers care about different units. Zero denominators per capability is what almost everyone has, and that is the actual problem. Do not let the search for the perfect single metric become another reason to not choose.

Three ways this goes wrong

The vanity denominator. Chosen because it trends nicely, quietly redefined at the point it stops. This is usually not dishonesty. It is someone genuinely improving a definition at a moment that happens to be convenient, which is why dating the definition matters so much.

Acceptance blindness. Measuring generated volume as though produced equals delivered. The tell is a program with excellent throughput numbers where nobody can tell you what proportion of output actually got used.

The unattributable metric. Technically sound, methodologically defensible, owned by nobody, discussed monthly, acted on never. This is the most common end state for a metrics program and it looks like success for about three quarters.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, make the Fox’s offer. Do not impose a denominator. Ask for one, accept whatever they propose, and hold them to it for two years. You will get a better metric than you would have chosen and considerably more commitment to it than you would have got by mandating one.

If you sit in the Crocodile’s chair, instrument acceptance, not just completion. This is the hardest signal in the stack and the most valuable, and if it is not designed in now it will require a retrofit later that nobody will fund.

If you sit in the Mandrill’s chair, define what accepted means for your use case before anybody builds anything. You will discover this takes longer to agree than you expect, and that the argument itself is worth having, because it surfaces that your own teams disagree about what good output looks like.

For everyone: write the denominator down and put a date on it. Undated definitions drift, and they take the trend line with them without anybody noticing until somebody builds a business case on eighteen months of two different measurements.

Jungle Lesson 6

Pick the denominator before you build the dashboard, and pick the one that counts what was accepted rather than what was produced. A capability that generates a thousand things nobody used is not efficient. It is fast at being wrong.

Next time: every territory in the jungle wants the Watering Hole to exist and not one of them wants it on their books. The Mandrill argues brilliantly and in bad faith, the Sloth arrives with a framework roughly six weeks after it would have been useful, and the question of who pays for the shared foundation turns out to be political rather than technical, which is why nobody solves it with a tool. Lesson 7 is about allocation, showback and chargeback.

If your organization has been discussing metrics for more than two quarters without choosing one, the problem is not that you lack a framework. You have several. The problem is that choosing means accepting a smaller number on purpose, and nobody wants to be the one who did that.

Scaling AI FinOps | Lesson 5: The Pilot Trap

Somebody finally counted the pilots.

It had not occurred to anyone to do this, because each one had been approved separately, by a different animal, at a different time, for a defensible reason. The count came out at a hundred and eighty something. The precise number was disputed for a while, on the grounds that several of them might not technically still be running, which turned out to be a question nobody could answer.

The Hyena found this genuinely funny for about four minutes, which is a long time for a hyena.

The Fox went through the list looking for the ones that had produced something. He found eleven. Of those eleven, four had been quietly switched off after the demonstration, two had become the departmental builds on the Tortoise’s list from the previous season, and the remaining five were still running as pilots, funded quarter to quarter, eighteen months after proving whatever they had been set up to prove.

“So what happened to the other hundred and seventy?” asked the Crow.

“Mostly nothing,” said the Fox. “They ran, they showed something, everyone was pleased, and then the person who cared moved on to something else.”

The Crocodile opened one eye.

“You didn’t buy a capability,” he said. “You bought a hobby, a hundred and eighty times.”

The arithmetic nobody does

Here is the calculation that nobody performs, because each pilot is approved in isolation and isolation is where this problem hides.

Every pilot carries a fixed cost that has nothing to do with its scope. Getting access sorted. Wiring up to a data source. Standing up somewhere to run it. Building enough evaluation to know whether it worked. A security review. A privacy review. The meetings, which are not free, and which involve expensive animals. Somebody writing a summary at the end.

That fixed cost does not shrink because the pilot is small. A two week experiment and a two month one pay roughly the same setup tax. Which means the true cost of a pilot is somewhere between two and five times what appears on its approval, and almost all of it is invisible because it is spread across other people’s time.

Now multiply by a hundred and eighty.

The number that comes out of that is almost always larger than whatever line item the organization is currently anxious about. In the jungle’s case it was substantially larger than the inference bill that had started this whole conversation four posts ago, and it had never once been discussed as a single number, because it had never once existed as a single number.

Why nothing compounds

The arithmetic is bad. The structural problem underneath it is worse.

In a healthy system, the fortieth thing you build is cheaper than the fourth, because the fourth left something behind. Some infrastructure, some patterns, some hard-won knowledge about what breaks. That is what makes a platform a platform.

Pilots leave nothing behind. Each one builds its own retrieval, its own evaluation, its own prompt scaffolding, its own access pattern, its own little arrangement with whoever owns the data. When it ends, all of that goes with it. The team disbands, the environment gets reclaimed, and the summary document goes into a folder that will be reorganized next year.

So pilot forty costs the same as pilot four. There is no curve. There is a straight line with a bad slope, extending as far as anyone is willing to keep funding it, and the organization experiences this as a series of individually reasonable decisions.

This is the thing to understand about the pilot trap. It is not that the pilots are bad. Several of them were excellent. It is that a hundred and eighty of them, run this way, produce exactly as much accumulated capability as one of them does, which is none.

The graduation problem

The five survivors on the Fox’s list deserve their own attention, because they illustrate a failure that is somehow both obvious and universal.

A pilot is funded to prove something. When it proves it, the funding logic has been satisfied, and the thing that made the money flow no longer applies. There is usually no defined path from proof to production, no budget line waiting, and no team whose job it is to receive it.

So the successful pilot ends up in exactly the same position as the failures. Still running, still on temporary funding, still argued about every quarter, and slowly starved by an organization that has moved on to approving the next batch.

I have watched genuinely good capabilities die of this. Not rejected. Not found wanting. Simply never picked up, because success was not something anybody had designed for. The Mandrill put it well when he finally understood what he was looking at: “We’ve built a system that can start things and cannot finish them, and we call the starting part innovation.”

The portfolio inversion

What good looks like is roughly the inverse of what most organizations have.

A small number of capabilities. A large number of use cases running on them. Not many initiatives each carrying its own complete stack, but a handful of well-built things that many parts of the business consume.

Almost every enterprise I have seen has this exactly backwards. Dozens or hundreds of initiatives, each with its own everything, and no shared foundation underneath any of it. Then somebody proposes consolidation, and the resistance is immediate, because every one of those initiatives has an owner who experiences consolidation as losing control of their thing.

The economics are not subtle. In the capability model, the fixed cost is paid once and amortized across everything that uses it. The marginal cost of the next use case falls as the platform matures. The cost curve bends. In the pilot model, the fixed cost is paid every single time, and the curve does not bend, ever, because there is nothing for it to bend around.

What this sets up

I am deliberately not going to resolve this here, because the resolution takes a whole act and doing it badly in three paragraphs is how this subject usually gets ruined.

What a capability actually is, how it gets funded when your organization’s entire financial machinery is built around things that end, who owns it, and what has to be traded to make it politically survivable, all of that is Act IV. It is the part of the series I expect to be most argued with.

For now the useful thing is to have the number. Count your pilots, multiply by an honest fixed cost, and look at the total. That single figure changes more conversations than any framework I could give you, because it converts a hundred and eighty reasonable individual decisions into one unreasonable aggregate one, which is what it was all along.

Three ways this goes wrong

Innovation theater by volume. The number of pilots becomes the metric that gets reported upward. Once that happens the incentive is to start things, not to finish them and certainly not to kill them, and the count grows because the count is the point.

The graduation cliff. No funding path from proof to production, so success is fatal. The tell is that your longest running pilots are also your best ones, which should be alarming and is usually described as pragmatism.

Capability in name only. A platform team is created, everyone agrees consolidation is the right idea, and then each territory carries on building its own stack on top of the shared one. You now pay the duplicated fixed costs and a governance layer, which is worse than where you started, and it is very common.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, get the count and multiply it by your honest fully loaded fixed cost per pilot. Include the meetings and the reviews. Bring the total to your next steering meeting as a single number with no commentary attached and let the room react to it.

If you sit in the Crocodile’s chair, list what every pilot has rebuilt from scratch. Retrieval, evaluation, access, deployment, monitoring. That list is not a complaint. It is the specification for your platform, written for you, by evidence.

If you sit in the Mandrill’s chair, stop starting things. Pick the three use cases in your territory that genuinely matter and kill the rest publicly, so that people believe the change is real. Quiet cancellation teaches nobody anything.

For everyone: ask what happens to a pilot that succeeds. If nobody can answer, you do not have a pipeline. You have a graveyard with excellent intentions and a quarterly approval process.

Jungle Lesson 5

Pilots are cheap individually and ruinous collectively, because each one pays the setup cost again and none of them make the next one easier. A hundred experiments is not a strategy with good coverage. It is the same experiment, funded a hundred times, by people who have not met.

That closes the first act. Five lessons in, the jungle now knows that its bill is shaped differently from anything it has managed before, that three groups have been arguing in three currencies, that the layer everyone budgeted for is not the layer that hurts, that the most interesting demand signal in the building was sitting on a compliance officer’s desk, and that a hundred and eighty reasonable decisions can add up to one unreasonable one.

What it still cannot do is answer the Crow’s original question. Are we getting enough for what we are paying? That requires learning to count, which is the second act, and it starts with the least glamorous and most consequential decision in the whole discipline: choosing what to divide by. Lesson 6 is about picking a denominator, and why choosing turns out to be considerably harder than complaining about not having chosen.

If you have never counted your pilots, the count is usually a bad afternoon and a very good quarter. I have yet to see an organization regret doing it, and I have yet to see one where the number was smaller than expected.

Scaling AI FinOps | Lesson 4: Termites in the Walls

The Tortoise arrived at the review with a list, which she put on the table without saying anything, and then waited.

Tortoises are the most underestimated animal in any organization. They are slow, they are armored, and they outlive absolutely everyone in the room, which means that over a long enough period they are usually the only creature who remembers what was decided and why. This one had been compiling her list for two seasons without mentioning it, because she had wanted it to be complete before anyone had the chance to argue with it.

It ran to four pages.

The Fox read it with an expression that changed twice. The first change was surprise at the length. The second was the recognition, about halfway down the second page, that three of the entries had been started by his own team, one of them by a person who reported directly to him.

“How long have you had this?” he asked.

“A while.”

“Why didn’t you bring it earlier?”

“Because,” said the Tortoise, “the last time somebody brought a list like this to a room like this, everything on it got banned within a fortnight, and eighteen months later there was a longer list that nobody had.”

Why this is not shadow IT

Every enterprise has had shadow IT and every enterprise thinks it knows how to handle this. It does not, because the two problems only look alike from a distance.

Shadow IT had friction. Somebody had to find a tool, pay for it, get it working, and usually get it past a network. There was a procurement event, an invoice, a login, an integration. There were artifacts, and artifacts are catchable.

Shadow AI has almost none of that. The entry cost is close to zero. There is frequently no infrastructure footprint at all. And in the largest category, which I will come to, there is no procurement event whatsoever, because the capability arrived inside software your organization already licensed and somebody simply switched it on.

There is, quite often, nothing to catch. Which means the entire enforcement instinct that your organization built over fifteen years is pointed at a problem with a different shape.

The three tiers

What the Tortoise’s list actually contained, once it was sorted, fell into three tiers.

Tier one: personal subscriptions. Individuals expensing tools, or in a fair number of cases not expensing them and paying out of pocket because it was easier than asking. Visible in expense data if anyone looks. Usually small in cost and disproportionately large in the anxiety it generates, because it is the tier everyone can see and therefore the tier everyone talks about.

Tier two: features inside software you already own. The assistant embedded in the document tool. The summarization inside the service platform. The drafting feature in the communications suite. Nobody procured these. Nobody decided to adopt them. A vendor shipped an update and several hundred animals started using something new on a Tuesday.

Tier three: departmental builds. A team with some technical capability builds something real on the sanctioned platform, and it never registers as a project because nobody filled in a form. Often genuinely good. Almost always without an owner named anywhere, a runbook, or a line in any budget.

The tier nobody sees

Tier two is the biggest and the least visible, and I think it is the most under-discussed problem in the whole field.

Your organization is already paying for it. The cost sits inside a bundled licence line that was negotiated for something else entirely, which means it does not appear in any AI cost report, is not attributed to any capability, and shows up in no conversation about adoption. From a financial reporting perspective it is not AI spend at all. It is last year’s software renewal.

Meanwhile several hundred people are putting real work into it daily. The Tortoise’s list had more entries in tier two than in the other two combined, and every one of them had been discovered by her opening an administrative console and reading, which took her an afternoon.

This is worth sitting with. The largest category of AI use in most organizations is one that costs nothing extra, appears in no report, and was never decided on by anybody.

What the exposure actually is

The reflex is to worry about cost. In my experience cost is rarely the real exposure here, and treating it as such gets you a program that solves the cheap problem.

The real exposure is threefold. Data leaving where it was supposed to stay. Retention you did not choose and cannot describe. And, most seriously, a business process that now depends on something with no owner, no runbook, no monitoring and no budget.

That third one is what should keep people awake. Somewhere in your organization right now there is a month end process, or a customer response path, or a regulatory submission, that has quietly come to depend on something a single person set up in an afternoon. That person may have since moved teams. When it breaks, nobody will know what it was, what it did, or how to restore it, and the discovery will happen at the worst possible moment because that is when dependencies reveal themselves.

The reframe

Here is the part that takes most leadership teams a while to accept.

Every entry on that list is somebody solving a real problem, on their own initiative, without being asked, and usually without being resourced. Not one of those animals woke up wanting to circumvent anything. They had work to do and something in front of them made it easier.

Which means the Tortoise’s four page list is the single best product roadmap in the building. It is demand, validated by the fact that people were willing to work around the organization to get it. No survey will ever produce information that good. No workshop will surface it. Your people have already told you what they need, in the most credible way available, which is by doing it.

The Fox, to his credit, worked this out about ten minutes into the meeting and stopped being defensive about the three entries belonging to his own team. “So this is the backlog,” he said. The Tortoise, who had been waiting two seasons for someone to say precisely that, allowed herself to look mildly pleased, which on a tortoise is a very small movement.

Amnesty, and what makes it work

The practical mechanism is an amnesty. A defined window during which declaring what you built gets you support, migration help and a proper budget line, rather than a conversation with someone about policy.

Three things make an amnesty work, and all three are required.

  • Credible non-punishment. If one person gets disciplined during the window, the window is over regardless of what the announcement said. This has to be stated at a level senior enough that people believe it.
  • A genuinely better sanctioned alternative. Amnesty with nothing to migrate to is just a census. People declare, nothing improves, and they carry on exactly as before with more paperwork and less trust.
  • A real deadline. Open-ended amnesties never close and therefore never create the moment of decision that makes them work.

Discovery, meanwhile, does not require surveillance and should not use it. Expense line analysis finds tier one. Reading your own administrative consoles finds tier two, and finds most of it in a single afternoon. Tier three is found by asking, which works considerably better than people expect once the amnesty is credible.

Three ways this goes wrong

The crackdown. A blanket ban, announced firmly. Usage does not stop. It moves to personal devices, personal accounts and personal networks, where you have no visibility at all. Your exposure has increased and your ability to see it has gone to zero, which is the exact opposite of what the ban was for.

Amnesty with nowhere to go. Everyone declares, the sanctioned platform is worse than what they were using, and within a quarter the list is growing again. You have spent your credibility and bought a snapshot.

Counting the wrong risk. A program obsessed with subscription costs, producing careful reports on a small number, while an unowned workflow sits inside a month end close and nobody has noticed.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, run an expense line analysis for AI subscriptions this month, then go and look at what AI features are bundled into your three largest software renewals. The second exercise will find more than the first and cost you nothing.

If you sit in the Crocodile’s chair, inventory what is already licensed and switched on. Most of tier two is sitting in administrative consoles you already have access to. This is an afternoon of work and it is the highest return afternoon available to you this quarter.

If you sit in the Mandrill’s chair, treat your territory’s list as a demand backlog rather than a compliance finding, and fund the top three properly. You will get better outcomes and considerably more goodwill than any enforcement action would have produced.

For everyone: for every instance you discover, record the problem it was solving in a column next to the cost. That column is worth more than the cost column, and almost nobody creates it.

Jungle Lesson 4

Shadow AI is not an act of defiance, it is an unfunded requirement with a credit card. Kill the tool and the requirement is still there, only now it is invisible to you and expensive to somebody else.

Next time: the jungle counts its pilots for the first time and the number is genuinely absurd. Every one of them was individually cheap and individually defensible. Together they consumed most of a year, and not one of them made the next one any easier to start. Lesson 5 closes the first act with the pilot trap.

If you have someone in your organization who has quietly been keeping a list like the Tortoise’s, find out. They are usually in risk or compliance, they have usually been waiting to be asked, and they are usually the best informed person in the building.

Scaling AI FinOps | Lesson 3: Anatomy of an AI Invoice

The Beaver offered to show the Crow around the Watering Hole, and she accepted, mostly because she had approved funding for it twice and had never actually seen it.

Beavers cannot hear running water without building something. This is not a character flaw, it is a compulsion, and it is the reason the Watering Hole worked at all. It is also the reason it had considerably more moving parts than anyone in the Canopy realized.

He started with the part she knew about. “This is where the thinking happens,” he said. “This is the bit everyone means when they say the word.”

“Right,” said the Crow. “That’s the line in my budget.”

“That’s about forty percent of it.”

She stopped walking.

He opened a panel she had not noticed. Behind it, something was running steadily, working through the jungle’s documents, converting them into a form the system could search. It had been running all night. It ran every night.

“What is that costing us?”

“Depends how much anyone wrote this week.”

“Not how much anyone used it?”

“No,” said the Beaver. “That one doesn’t care whether anybody uses it.”

The Crow, who had spent a career learning that costs follow activity, took a moment with that. Then she asked him to open every other panel.

The layer you budgeted

Almost every organization budgets for inference and is surprised by everything else. Not because the other layers are hidden, exactly, but because the mental model arrived from somewhere else. You buy a thinking machine, you pay per thought. That is the shape of the thing as it is described in every article, every business case, and every conversation with a vendor.

The shape is wrong. Inference is usually somewhere between a third and half of the true cost of a running capability, and the proportion falls as the capability matures, because the supporting layers keep growing while the thinking gets cheaper.

Here is the full stack, in roughly the order it surprises people.

The seven layers

  • Inference. The one you budgeted. Scales with usage, drops in unit price over time, and is the only layer most organizations can report on.
  • Retrieval and indexing. Embedding your content, storing it, and re-processing it every time it changes. This scales with the size of your corpus and the rate at which people write things, not with how much anyone uses the system.
  • Orchestration and glue. Queues, state, workflow, the compute that coordinates rather than thinks. Individually cheap, structurally permanent, and it never appears on a slide.
  • Evaluation. Running your test sets, regression checks and scoring. Scales with release frequency, which means it grows precisely as the team gets better at shipping.
  • Observability and logging. Capturing what went in and what came out, at volume, and keeping it for as long as somebody with an armored shell says you have to.
  • Human in the loop. Review, correction, escalation, the person who checks the output before it goes anywhere consequential. This is a labor cost and it usually lands in a completely different budget.
  • Rework. Output that was wrong, redone, or quietly discarded. Costs you twice and appears nowhere at all.

Seven layers. Most organizations can see one of them clearly, two of them vaguely, and are actively surprised by the rest.

The two that grow when nobody is using it

Retrieval and evaluation deserve their own paragraph, because they break the intuition that cost follows activity, and that intuition is the load-bearing wall of every cost conversation your finance team has ever had.

Retrieval cost tracks your content, not your usage. If a content team has a productive quarter and publishes four hundred new documents, your retrieval bill goes up, and not one additional animal has asked the system a single question. If somebody reorganizes a document library, everything gets reprocessed. If a well-intentioned team decides to index a large archive that nobody has opened since the last reorganization, you will pay to keep that archive current forever.

Evaluation cost tracks your release cadence. A team that ships weekly runs its test suites weekly. As the suites grow, which they should, the cost of every release grows with them. This is the correct behavior and it is still a cost line that nobody forecast, and it arrives at exactly the moment when a program is going well, which is the worst moment for a finance surprise.

Neither of these responds to a usage cap. Neither of these shows up in a conversation about adoption. Both of them appear on the invoice.

The layer in someone else’s budget

Human review is the one I see cause the most damage, and it is not close.

The capability produces output. Somebody checks it. That checking is real work performed by real people, and in almost every organization I have seen, those people sit in an operations budget, a service budget, or a shared function that has nothing to do with the AI program.

So the capability looks cheap. Its cost line contains inference and some infrastructure. The review labor it generates is invisible to it, sits somewhere else, and grows in direct proportion to how much the capability is used.

Then somebody decides to scale it, on the basis of a unit cost that excluded most of the marginal cost, and the operations budget quietly absorbs the difference until it cannot, which is usually about two quarters later and always somebody else’s problem to explain.

The Mandrill, when this was pointed out to him, did something I have a lot of respect for. He asked for the review hours to be attributed to his capability, knowing it would make his own business case look worse. It did make it look worse. It also made it real, and the version of the decision he made afterwards was a better one.

The invisible layer

Rework is the seventh layer and the one nobody can measure, which is precisely why it is worth naming.

Output that was wrong and had to be redone. Output that was fine and got discarded because the person did not trust it. Attempts that succeeded on the third try. Work generated, reviewed, rejected, and generated again.

You cannot instrument this precisely. What you can do is stop pretending it is zero, because zero is the number currently in every business case in your organization. A documented estimate, even a rough one, is a better input than an implicit assumption that the failure rate is nil. And the act of estimating it tends to surface how little anybody knows about acceptance rates, which is a useful thing to discover in a planning meeting rather than in an audit.

Fixed and variable

The last thing worth pulling out of the stack is which parts are fixed and which are variable, because that ratio determines the shape of every architecture decision you will make later.

Inference is variable. Human review is variable. Retrieval and evaluation are mostly fixed with respect to usage and variable with respect to other things entirely. Orchestration and observability are largely fixed once built.

A capability with a high fixed proportion has terrible unit economics at low volume and excellent unit economics at high volume. A capability that is almost entirely variable has flat unit economics forever. Neither is better. But if you do not know which one you have, you cannot answer the question of whether scaling it will help or hurt, and that question is going to come up.

One thing the Beaver said as we finished the tour has stayed with me. “Everyone asks me how much a question costs,” he said. “Nobody asks me how much it costs to be ready for one.”

Three ways this goes wrong

Inference tunnel vision. Months spent optimizing the visible layer while the other six grow unwatched. The optimization work is real and the savings are real and the total bill goes up anyway, which does considerable damage to the credibility of everyone involved.

Corpus creep. Retrieval costs rising steadily with no change in usage, and nobody connecting it to the fact that three teams have been busy publishing. This one can run for a year before anybody asks the right question.

The orphan labor line. Human review costs sitting in an operations budget, so the capability appears far cheaper than it is, and gets scaled on a unit cost that was never true.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, pick one capability and demand the cost broken into all seven layers. Just one. You do not need the whole estate to make the point, and the exercise of producing it will teach the team more than the answer teaches you.

If you sit in the Crocodile’s chair, put retrieval and evaluation onto their own cost lines today, separately from inference. They are the two that grow silently and they are invisible inside an aggregate.

If you sit in the Mandrill’s chair, find the human review hours your capability generates and insist they are attributed to it. Yes, it makes your numbers worse. It also means the decision you make next is based on something true, and you will be the only person in the room who can say that.

For everyone: put a number on rework, even a guess, and write down how you arrived at it. Any documented estimate beats the zero that is currently sitting there by default.

Jungle Lesson 3

Inference is the layer you budgeted and rarely the layer that hurts. The expensive parts of an AI capability are the ones that grow when nobody is using it, and the ones that get paid for out of somebody else’s budget.

Next time: the Tortoise puts a list on the table. She has been quietly compiling it for two seasons and it is considerably longer than anyone expected, and several entries turn out to have been started by the Fox’s own team. Lesson 4 is about shadow AI, why the instinct to stamp it out is the most expensive available response, and why that list is the best product roadmap in the building.

If anyone has ever shown you a full seven layer breakdown of a live capability without being asked, keep them. They are rarer than they should be.

Scaling AI FinOps | Lesson 2: Three Tribes, One Bill

The Fox explained the number three times on the same day, to three different audiences, and gave three completely different answers. All three were true. That was the problem.

At nine in the morning he told the Crow that the overage came from higher than forecast request volume in two territories, and that the average unit cost had actually improved quarter on quarter.

At eleven he told the Crocodile’s team that the retrieval layer was over-fetching, that context sizes had grown by roughly half since launch, and that nobody had touched the routing configuration since the pilot.

At three he told the Mandrill that adoption was running ahead of plan, that his territory was seeing measurably faster proposal turnaround, and that the spend reflected genuine business pull.

Every one of those statements was accurate. Not one of them could be reconciled with the other two. And by Friday all three animals were quietly certain that the Fox was telling each of them what they wanted to hear.

The Hyena had been listening from a branch for most of the day, because it was more entertaining than working. She offered the only useful observation anyone made.

“He isn’t lying to you,” she said. “He’s answering in three currencies and none of you can do the exchange rate.”

The meeting is not the problem

I have watched a lot of organizations conclude that their AI cost conversations are failing because of politics. Finance is being difficult. Engineering is being evasive. The business is being unrealistic. Somebody suggests a workshop.

It is almost never politics. Politics is what grows in the gap afterwards, once three competent groups have spent a quarter failing to understand each other and have each drawn the obvious conclusion about the other two.

The underlying failure is more boring and much more fixable. Three groups are describing the same system in three vocabularies, each internally coherent, none of which converts into the others. Nobody has defined what the shared thing being counted actually is. So every meeting is a currency exchange conducted by people who each believe they are already speaking the common language.

Three currencies, no exchange rate

The Crow thinks in periods, budget lines and variance. Her unit is a committed number and her time horizon is the fiscal year. When she asks what something costs, she means what will appear against a line she has to defend, in a period she has to close. This is not narrow-mindedness. It is the job, and the job has a shape.

The Crocodile thinks in requests, throughput, utilization and latency. His unit is a system under load and his time horizon is roughly now. When he says the cost is fine, he means the cost per request is trending down, which it genuinely is, and which has nothing whatsoever to do with the Crow’s question.

The Mandrill thinks in outcomes, cycle time and headcount. His unit is a business result and his time horizon is whatever he committed to at the last review. When he says it is working, he means his territory is producing more of something he cares about, which is also true, and which nobody has connected to either of the other two answers.

Each vocabulary is correct. Each one is load-bearing for the animal using it. And each one answers a question the others were not asking.

What makes this worse than a normal translation problem is that all three groups have been in enough meetings together to believe they have already adjusted. The Crow has learned to say request. The Crocodile has learned to say budget. They are using each other’s words for their own concepts, which is not translation. It is a false cognate, and false cognates are more dangerous than an unknown language, because nobody notices the error.

Why dashboards do not fix this

The standard response at this point is to build something. Usually a dashboard, occasionally a whole platform, and there is a business case involving single source of truth.

Watch what actually gets built. Engineering data, rendered attractively, placed in front of a finance audience. Requests per day. Tokens by team. Average latency. Cost per thousand calls, in a nice color.

That is not translation. That is subtitling. The words are now legible to the Crow and the meaning is not, because the underlying unit never changed. She can read every number on the screen and still cannot answer the only question she has, which is whether this was worth it. So she asks that question again, out loud, in the review, and the room concludes that finance does not understand the technology.

Finance understands the technology fine. Finance is being handed the numerator and asked to reason about a ratio.

The Owl, at this point, produced a diagram. It mapped all three vocabularies onto a single canvas with color-coded arrows showing the relationships between them. It was, I want to be fair here, completely accurate. Everyone agreed it was very clear. Nobody’s behavior changed by a single decision, because a map of a disagreement is not a resolution of it, and the Owl had drawn the territory rather than deciding what to call the things in it.

The shared unit of account

What is missing is a shared unit of account. One thing that all three tribes agree describes the same object, sitting deliberately between the technical metric and the business outcome, close enough to each that both can reach it.

Not tokens. That is engineering’s currency and it belongs in engineering’s meetings, where it is genuinely useful. Not annual budget. That is the Crow’s currency and it aggregates away everything actionable. Not revenue. That is the Mandrill’s and it has too many other parents to attribute cleanly.

Something in the middle. A resolved case. An accepted draft. A processed document. A completed review. The unit of work the organization actually performs, which happens to be the one thing all three animals were already talking about without noticing.

Four tests for whether you have picked a good one.

  • Countable without new instrumentation. If measuring it requires a project, you will not measure it, and the definition will quietly become an estimate within two quarters.
  • Meaningful without explanation. If the Mandrill needs a preamble to understand what it represents, it will not survive contact with a steering committee.
  • Attributable to someone. An excellent metric that belongs to nobody produces no decisions. It produces observations, which is a different and much less useful thing.
  • Stable across quarters. If the definition moves, the trend line is fiction, and you will not notice until someone builds a business case on it.

The conversion chain

Once you have the unit, the argument becomes tractable, because you can lay out the chain and see where it is solid and where it is assumed.

Technical unit to work unit to business outcome to money. Four links. Requests to resolved cases. Resolved cases to reduced backlog. Reduced backlog to something in the ledger.

Most organizations attempt to jump from the first link to the last in a single move, and lose the argument at the first joint, because the person on the other side can feel the gap even if they cannot name it. The Crow could not have told you which link in the Fox’s Tuesday argument was weak. She could tell you immediately that something was.

The discipline is to draw all four links explicitly and then mark honestly which ones are measured and which ones are assumed. Almost always the first link is measured, the last is asserted, and the two in the middle are where the real work sits. That is not an embarrassing finding. That is the map of what to go and do next, and it is worth more than any dashboard you could commission this quarter.

The Crocodile, who had been silent for most of this, opened one eye. “So we’ve spent six weeks arguing,” he said, “about a conversion nobody had written down.”

Three ways this goes wrong

Subtitling. The same engineering data in a prettier chart, refreshed monthly, satisfying nobody. The tell is that the review meeting is the same length as before and produces the same number of decisions, which is none.

The composite index. Someone, usually with good intentions and a background in analytics, builds a weighted score combining nine metrics into a single number between zero and a hundred. It goes up. Nobody can act on it, because no single lever moves it and no team owns it. The composite index is what organizations build instead of choosing, and choosing was the whole task.

Unit drift. The definition changes quietly between quarters, usually because someone improved the measurement, and the trend line becomes a comparison between two different things. This one is insidious because it happens for good reasons, and it is only discoverable if the definition was dated when it was written.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask each of the three groups to describe the same use case in one written sentence, separately, without conferring. Put the three sentences side by side. The gap between them is not a communication problem to be smoothed over in a workshop. It is the actual problem, now visible, in about twenty minutes.

If you sit in the Crocodile’s chair, publish the conversion chain for one capability, all four links, and mark clearly which links are measured and which are assumed. Do not wait until it looks good. The version with honest gaps in it is far more valuable than the version that took a quarter to make presentable.

If you sit in the Mandrill’s chair, own the business outcome end of the chain and define it in writing. Nobody else can do this and if you leave it vacant, somebody will cheerfully invent an outcome on your behalf, and you will be held to it.

For everyone: agree the unit of account in writing, with a date on it, before anyone commissions a dashboard. Tooling built on the wrong denominator is worse than no tooling, because it makes the wrong answer look rigorous, and rigorous wrong answers are much harder to dislodge than obvious ones.

Jungle Lesson 2

A disagreement you cannot resolve is usually a translation failure wearing a disagreement costume. Before you argue about the number, agree what the number is counting, write the definition down, and put a date on it.

Next time: the Beaver takes the Crow on a tour of the platform and keeps opening panels she did not know were there. It turns out that the layer everyone budgeted for is rarely the layer that hurts, and two of the expensive ones grow steadily whether anybody uses the thing or not. Lesson 3 is the anatomy of an AI invoice.

If your organization is currently three groups talking past each other about the same bill, I would be curious which of the three vocabularies is winning. In my experience it is whichever one the loudest person speaks.