Scaling AI FinOps | Lesson 2: Three Tribes, One Bill

The Fox explained the number three times on the same day, to three different audiences, and gave three completely different answers. All three were true. That was the problem.

At nine in the morning he told the Crow that the overage came from higher than forecast request volume in two territories, and that the average unit cost had actually improved quarter on quarter.

At eleven he told the Crocodile’s team that the retrieval layer was over-fetching, that context sizes had grown by roughly half since launch, and that nobody had touched the routing configuration since the pilot.

At three he told the Mandrill that adoption was running ahead of plan, that his territory was seeing measurably faster proposal turnaround, and that the spend reflected genuine business pull.

Every one of those statements was accurate. Not one of them could be reconciled with the other two. And by Friday all three animals were quietly certain that the Fox was telling each of them what they wanted to hear.

The Hyena had been listening from a branch for most of the day, because it was more entertaining than working. She offered the only useful observation anyone made.

“He isn’t lying to you,” she said. “He’s answering in three currencies and none of you can do the exchange rate.”

The meeting is not the problem

I have watched a lot of organizations conclude that their AI cost conversations are failing because of politics. Finance is being difficult. Engineering is being evasive. The business is being unrealistic. Somebody suggests a workshop.

It is almost never politics. Politics is what grows in the gap afterwards, once three competent groups have spent a quarter failing to understand each other and have each drawn the obvious conclusion about the other two.

The underlying failure is more boring and much more fixable. Three groups are describing the same system in three vocabularies, each internally coherent, none of which converts into the others. Nobody has defined what the shared thing being counted actually is. So every meeting is a currency exchange conducted by people who each believe they are already speaking the common language.

Three currencies, no exchange rate

The Crow thinks in periods, budget lines and variance. Her unit is a committed number and her time horizon is the fiscal year. When she asks what something costs, she means what will appear against a line she has to defend, in a period she has to close. This is not narrow-mindedness. It is the job, and the job has a shape.

The Crocodile thinks in requests, throughput, utilization and latency. His unit is a system under load and his time horizon is roughly now. When he says the cost is fine, he means the cost per request is trending down, which it genuinely is, and which has nothing whatsoever to do with the Crow’s question.

The Mandrill thinks in outcomes, cycle time and headcount. His unit is a business result and his time horizon is whatever he committed to at the last review. When he says it is working, he means his territory is producing more of something he cares about, which is also true, and which nobody has connected to either of the other two answers.

Each vocabulary is correct. Each one is load-bearing for the animal using it. And each one answers a question the others were not asking.

What makes this worse than a normal translation problem is that all three groups have been in enough meetings together to believe they have already adjusted. The Crow has learned to say request. The Crocodile has learned to say budget. They are using each other’s words for their own concepts, which is not translation. It is a false cognate, and false cognates are more dangerous than an unknown language, because nobody notices the error.

Why dashboards do not fix this

The standard response at this point is to build something. Usually a dashboard, occasionally a whole platform, and there is a business case involving single source of truth.

Watch what actually gets built. Engineering data, rendered attractively, placed in front of a finance audience. Requests per day. Tokens by team. Average latency. Cost per thousand calls, in a nice color.

That is not translation. That is subtitling. The words are now legible to the Crow and the meaning is not, because the underlying unit never changed. She can read every number on the screen and still cannot answer the only question she has, which is whether this was worth it. So she asks that question again, out loud, in the review, and the room concludes that finance does not understand the technology.

Finance understands the technology fine. Finance is being handed the numerator and asked to reason about a ratio.

The Owl, at this point, produced a diagram. It mapped all three vocabularies onto a single canvas with color-coded arrows showing the relationships between them. It was, I want to be fair here, completely accurate. Everyone agreed it was very clear. Nobody’s behavior changed by a single decision, because a map of a disagreement is not a resolution of it, and the Owl had drawn the territory rather than deciding what to call the things in it.

The shared unit of account

What is missing is a shared unit of account. One thing that all three tribes agree describes the same object, sitting deliberately between the technical metric and the business outcome, close enough to each that both can reach it.

Not tokens. That is engineering’s currency and it belongs in engineering’s meetings, where it is genuinely useful. Not annual budget. That is the Crow’s currency and it aggregates away everything actionable. Not revenue. That is the Mandrill’s and it has too many other parents to attribute cleanly.

Something in the middle. A resolved case. An accepted draft. A processed document. A completed review. The unit of work the organization actually performs, which happens to be the one thing all three animals were already talking about without noticing.

Four tests for whether you have picked a good one.

  • Countable without new instrumentation. If measuring it requires a project, you will not measure it, and the definition will quietly become an estimate within two quarters.
  • Meaningful without explanation. If the Mandrill needs a preamble to understand what it represents, it will not survive contact with a steering committee.
  • Attributable to someone. An excellent metric that belongs to nobody produces no decisions. It produces observations, which is a different and much less useful thing.
  • Stable across quarters. If the definition moves, the trend line is fiction, and you will not notice until someone builds a business case on it.

The conversion chain

Once you have the unit, the argument becomes tractable, because you can lay out the chain and see where it is solid and where it is assumed.

Technical unit to work unit to business outcome to money. Four links. Requests to resolved cases. Resolved cases to reduced backlog. Reduced backlog to something in the ledger.

Most organizations attempt to jump from the first link to the last in a single move, and lose the argument at the first joint, because the person on the other side can feel the gap even if they cannot name it. The Crow could not have told you which link in the Fox’s Tuesday argument was weak. She could tell you immediately that something was.

The discipline is to draw all four links explicitly and then mark honestly which ones are measured and which ones are assumed. Almost always the first link is measured, the last is asserted, and the two in the middle are where the real work sits. That is not an embarrassing finding. That is the map of what to go and do next, and it is worth more than any dashboard you could commission this quarter.

The Crocodile, who had been silent for most of this, opened one eye. “So we’ve spent six weeks arguing,” he said, “about a conversion nobody had written down.”

Three ways this goes wrong

Subtitling. The same engineering data in a prettier chart, refreshed monthly, satisfying nobody. The tell is that the review meeting is the same length as before and produces the same number of decisions, which is none.

The composite index. Someone, usually with good intentions and a background in analytics, builds a weighted score combining nine metrics into a single number between zero and a hundred. It goes up. Nobody can act on it, because no single lever moves it and no team owns it. The composite index is what organizations build instead of choosing, and choosing was the whole task.

Unit drift. The definition changes quietly between quarters, usually because someone improved the measurement, and the trend line becomes a comparison between two different things. This one is insidious because it happens for good reasons, and it is only discoverable if the definition was dated when it was written.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask each of the three groups to describe the same use case in one written sentence, separately, without conferring. Put the three sentences side by side. The gap between them is not a communication problem to be smoothed over in a workshop. It is the actual problem, now visible, in about twenty minutes.

If you sit in the Crocodile’s chair, publish the conversion chain for one capability, all four links, and mark clearly which links are measured and which are assumed. Do not wait until it looks good. The version with honest gaps in it is far more valuable than the version that took a quarter to make presentable.

If you sit in the Mandrill’s chair, own the business outcome end of the chain and define it in writing. Nobody else can do this and if you leave it vacant, somebody will cheerfully invent an outcome on your behalf, and you will be held to it.

For everyone: agree the unit of account in writing, with a date on it, before anyone commissions a dashboard. Tooling built on the wrong denominator is worse than no tooling, because it makes the wrong answer look rigorous, and rigorous wrong answers are much harder to dislodge than obvious ones.

Jungle Lesson 2

A disagreement you cannot resolve is usually a translation failure wearing a disagreement costume. Before you argue about the number, agree what the number is counting, write the definition down, and put a date on it.

Next time: the Beaver takes the Crow on a tour of the platform and keeps opening panels she did not know were there. It turns out that the layer everyone budgeted for is rarely the layer that hurts, and two of the expensive ones grow steadily whether anybody uses the thing or not. Lesson 3 is the anatomy of an AI invoice.

If your organization is currently three groups talking past each other about the same bill, I would be curious which of the three vocabularies is winning. In my experience it is whichever one the loudest person speaks.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.