Scaling AI FinOps | Lesson 6: Choosing a Denominator

The rains came, which in the jungle means budget season, and the Crow did something the Fox had not been expecting.

She offered him a deal.

“Pick your denominator,” she said. “Whatever unit of work you think this capability produces. You choose it, not me. I will fund against it and I will not argue with the one you pick.”

The Fox, who had spent two seasons being told that finance did not understand what he was building, waited for the catch.

“You own it for two years,” said the Crow. “You do not get to change it when it stops flattering you.”

That was the catch. It was also, though it took him most of a month to see it, the fairest offer anybody had made him since this began.

He had assumed choosing would take an afternoon. It took three weeks, and he abandoned four candidates before he found one he was willing to sign his name to for two years. Somewhere in the middle of the second week he came to a realization that I think most people in his position eventually arrive at, usually later than he did.

Complaining about not having a denominator is enormously easier than choosing one, and he had been doing the easy version for a year.

Why choosing is hard

The four tests from Lesson 2 still apply. Countable without new instrumentation, meaningful without explanation, attributable to someone, stable across quarters. Those get you a shortlist.

What they do not tell you is the thing that makes the decision genuinely difficult, which is that every honest denominator makes your capability look worse than the dishonest one sitting next to it. Choosing well is an act of deliberately taking a smaller number, on purpose, and defending it. That is why so few organizations do it, and it is not because they lack a framework.

The gap between generated and accepted

Here is the distinction that does most of the work in this post.

Cost per generated output is easy to measure and almost useless. Every system produces output. Producing output is the one thing you can absolutely rely on it to do, whether or not the output was any good, whether or not anyone read it, whether or not it went anywhere.

Cost per accepted output is hard to measure and true. Accepted meaning used. Sent. Merged. Approved. Acted upon. The thing where a person looked at what came out and decided it was good enough to carry forward.

Almost the entire value question lives in the gap between those two numbers, and organizations that measure the easy one systematically overstate their own performance, sometimes by a lot. I have seen acceptance rates that would have changed a funding decision if anyone had thought to look, sitting quietly inside a program reporting excellent volume metrics every month.

The unpleasant part is that acceptance is often the one thing nobody instrumented, because at pilot stage it did not matter. Everything was accepted at pilot stage. That was rather the point of a pilot, and it is why pilot metrics travel so badly into production.

Four archetypes

What this looks like in practice, with the acceptance definition attached, because the definition is the hard half.

  • Service and support work. Cost per resolved case. Accepted means the case closed and did not reopen within some window you choose in advance and then do not adjust.
  • Content and communication. Cost per accepted draft. Accepted means it went out, or went to the next stage, with or without editing. If you want to be rigorous, track the edit distance too, because a draft that gets rewritten entirely was not accepted, it was raw material.
  • Operations and processing. Cost per processed document. Accepted means it cleared without human correction. This is the cleanest of the four and the one most organizations can measure today if they look.
  • Engineering work. Cost per merged change. Accepted means it went into the main line. Suggestions that were dismissed are not free and should not be invisible.

Notice that in every case the acceptance definition contains a judgment call, and the judgment call belongs to the business rather than to engineering. This is why the Mandrill has to be in the room for this conversation, and why it goes badly when he is not.

The two traps

Two denominators are chosen constantly and are wrong in ways worth being explicit about.

Cost per user. This is the most common choice and the most reliably damaging one, because it rewards exactly the wrong behavior. Under cost per user, the way to improve your metric is to add more users. Adding users who barely touch the thing improves your number considerably. Meanwhile your heaviest users, the ones actually generating the value, make the metric look worse, which creates a quiet institutional pressure to discourage the people you most want to encourage. It is the metric that most reliably produces the wrong decision, and it is on more slides than any other.

Cost per token, or per request, or per call. This is a genuinely useful efficiency metric and it belongs to the Crocodile, in his own meetings, where it will do good work. It is not a business metric and it should not be near a steering committee, because it tells you how efficiently you are doing something without telling you whether the something was worth doing. A system can get steadily cheaper per request while producing steadily less value, and cost per request will report that as a success story every single month.

Does it survive a bad quarter

The test I would add to the four from Lesson 2 is this one, and it is the one that catches the metrics chosen for the wrong reasons.

Imagine the capability has a bad quarter. Volume down, quality complaints up, a territory unhappy. Does your denominator show that? Or does it stay flat, or improve, because of some property of how it is constructed?

If a metric only looks sensible when things are going well, it is not a measurement. It is a narrative device, and the moment it is needed it will not be there. The Fox discarded two of his four candidates on precisely this test, which is why the three weeks were well spent.

One more thing worth saying. More than one denominator per capability is fine, and often better, because different consumers care about different units. Zero denominators per capability is what almost everyone has, and that is the actual problem. Do not let the search for the perfect single metric become another reason to not choose.

Three ways this goes wrong

The vanity denominator. Chosen because it trends nicely, quietly redefined at the point it stops. This is usually not dishonesty. It is someone genuinely improving a definition at a moment that happens to be convenient, which is why dating the definition matters so much.

Acceptance blindness. Measuring generated volume as though produced equals delivered. The tell is a program with excellent throughput numbers where nobody can tell you what proportion of output actually got used.

The unattributable metric. Technically sound, methodologically defensible, owned by nobody, discussed monthly, acted on never. This is the most common end state for a metrics program and it looks like success for about three quarters.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, make the Fox’s offer. Do not impose a denominator. Ask for one, accept whatever they propose, and hold them to it for two years. You will get a better metric than you would have chosen and considerably more commitment to it than you would have got by mandating one.

If you sit in the Crocodile’s chair, instrument acceptance, not just completion. This is the hardest signal in the stack and the most valuable, and if it is not designed in now it will require a retrofit later that nobody will fund.

If you sit in the Mandrill’s chair, define what accepted means for your use case before anybody builds anything. You will discover this takes longer to agree than you expect, and that the argument itself is worth having, because it surfaces that your own teams disagree about what good output looks like.

For everyone: write the denominator down and put a date on it. Undated definitions drift, and they take the trend line with them without anybody noticing until somebody builds a business case on eighteen months of two different measurements.

Jungle Lesson 6

Pick the denominator before you build the dashboard, and pick the one that counts what was accepted rather than what was produced. A capability that generates a thousand things nobody used is not efficient. It is fast at being wrong.

Next time: every territory in the jungle wants the Watering Hole to exist and not one of them wants it on their books. The Mandrill argues brilliantly and in bad faith, the Sloth arrives with a framework roughly six weeks after it would have been useful, and the question of who pays for the shared foundation turns out to be political rather than technical, which is why nobody solves it with a tool. Lesson 7 is about allocation, showback and chargeback.

If your organization has been discussing metrics for more than two quarters without choosing one, the problem is not that you lack a framework. You have several. The problem is that choosing means accepting a smaller number on purpose, and nobody wants to be the one who did that.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.