Scaling AI FinOps | Lesson 24: Growing the Troop

The Mandrill wanted to hire a Head of AI FinOps.

It was, on every visible measure, the right call. The discipline had proved itself over two years. It had a rhythm, a ledger, an allocation model and a set of controls. It clearly needed an owner. The obvious next step for anything that works in an enterprise is to give it a leader, a team, a budget and a reporting line, and the Mandrill had reached that conclusion the way any competent executive would.

The Crocodile talked him out of it in about six sentences.

“I have watched this happen four times,” he said. “Capacity planning. Quality. Security. Cloud cost. Each one worked when it was distributed, so somebody centralized it, and each one turned into a function that produced excellent reports about decisions it was not in the room for. Then it got cut, and everyone said the discipline had not delivered. The discipline delivered. It was moved somewhere it could not reach anything.”

The Mandrill thought about it for a moment and changed his mind, out loud, in front of the room, which is the most senior thing anybody does in this entire story and considerably rarer than the reorganization the Fox proposed eight lessons earlier.

Why centralizing kills it

The mechanism is simple and it is the same every time.

The decisions that move the number are made by engineers choosing routing and context structure, and by product owners choosing scope and sufficiency. Those are the levers. Everything in Act III and half of Act IV lives in the hands of people doing the work.

A central team owns reports. It can describe what happened. It can identify where the money went. It cannot change a routing rule, cannot define an acceptance criterion, cannot decide a quality floor and cannot say no to a use case.

So you get a function with total visibility and no authority, producing increasingly sophisticated analysis of decisions being made elsewhere by people who have not read it. That is not a failure of the people in the function. They are usually good and they are usually frustrated. It is a structural mismatch between where the information sits and where the levers are.

The distributed model

A small central function and embedded accountability everywhere else.

The center owns standards, the ledger, shared tooling and arbitration. It defines how measurement works so that numbers are comparable across the estate, curates the ledger so it stays honest, and resolves disputes that teams cannot resolve themselves.

The capability teams own the decisions. Routing, context, scope, sufficiency, denominators, acceptance definitions, everything that actually changes the number.

Center sets the how of measurement. Teams own the what of decisions. If your center is making decisions, it is too big. If your teams cannot compare their numbers to each other, it is too small.

Four capabilities, not four roles

These have to exist somewhere in your organization. Deliberately described as capabilities rather than headcount, because the fastest way to get this wrong is to read them as a hiring plan.

Unit economics literacy in engineering. The people choosing architecture need to understand the cost consequence at the moment they are choosing, not in a review three weeks later. An engineer who knows what a cache miss costs makes a different decision from one who does not, and no governance process substitutes for that.

Technical literacy in finance. The Crow needs to know what a cascade is. Not to design one. To ask about it. A finance function that can ask what proportion of traffic goes to the top tier is worth more than any dashboard that would have answered the question unasked.

Value definition in the business. Denominators and acceptance criteria. This cannot be delegated to anyone else and it will be invented on your behalf if you leave it vacant.

Arbitration somewhere neutral. Somebody has to resolve cost against quality without owning either side. This is the one role that genuinely does need to sit in the center, and it needs to be senior enough that its decisions stick.

Literacy is the actual lever

The most useful thing in this post, and it is cheap, which is why it gets skipped in favor of something expensive.

Most organizations buy tooling to compensate for a literacy gap. The reasoning is reasonable: our finance team cannot interrogate this, so let us buy a platform that presents it in terms they understand.

What actually happens is that you have bought something that answers questions nobody knows how to ask. The platform is fine. The dashboards are good. Nobody in the room has the vocabulary to look at them and know which follow-up question would be devastating.

A finance team that understands the cost stack from Lesson 3 and the routing question from Lesson 12 asks better questions than any tool answers. That literacy costs a few days of somebody’s time and it compounds permanently. The tool costs considerably more and depreciates.

Invest in the literacy first. If you still need the tool afterwards, you will at least know what you need it for, which is a much better position from which to buy anything.

Roles to be skeptical of

Three job descriptions that get written and should usually not be filled as written.

The AI FinOps Analyst who only reports. Produces the monthly pack, attends the reviews, changes nothing. Frequently a capable person who will leave within eighteen months because they can see the problem more clearly than anyone.

The Cost Optimization Team with no authority over architecture. Recommends. Does not decide. Watches its recommendations get deprioritized against feature work every quarter, forever.

The Center of Excellence that excellent people leave. A perfectly good idea that becomes a holding pattern, because sitting adjacent to the work is less interesting than doing it, and the best people notice that first.

Career design

The structural point, which mirrors Lesson 22 exactly.

Nobody’s career currently rewards this work. There is no promotion attached to a cost per unit that fell. There is no recognition for the routing change that saved a great deal of money and produced no visible output.

So put it in objectives. Cost per unit as a named responsibility in engineering objectives. Denominator ownership in product objectives. Retirements alongside launches in portfolio objectives.

Until it is somebody’s actual job, described in the document that determines their progression, it will remain everybody’s concern and nobody’s work, and it will be done in the gaps by whoever happens to care. That works for about a year and it does not survive that person changing roles.

Three ways this goes wrong

The reporting function. Central team, beautiful outputs, no authority, no change, quietly disbanded in year three with the conclusion that the discipline did not deliver.

Tooling as a substitute for literacy. An expensive platform bought to answer questions the organization has not learned to ask, which it will therefore continue not to ask, at a higher cost.

Accountability without authority. Somebody named responsible for cost who cannot influence architecture, scope or funding. This is the most demoralizing job in enterprise technology and it is created constantly with the best of intentions.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, invest in technical literacy in your own team before you buy anything. Two days of your finance team learning the cost stack will change more meetings than any platform you could procure this year.

If you sit in the Crocodile’s chair, put cost per unit into engineering objectives. It changes behavior faster than any review forum, because it changes what people think about while they are making the decision rather than afterwards.

If you sit in the Mandrill’s chair, resist building the function. Distribute the accountability, keep the center small and standards-focused, and give it arbitration authority rather than delivery responsibility.

For everyone: name who arbitrates cost against quality. If the answer is nobody, then the loudest person in the room is arbitrating, and they are not doing it well because it is not their job and they do not know they are doing it.

Jungle Lesson 24

The decisions that move the number are made by the animals choosing the architecture and the animals choosing the scope, so a central team that owns the reports and none of the decisions will produce beautiful reports and no change. Keep the center small, distribute the accountability, and spend the money on literacy rather than on tooling that answers questions nobody knows how to ask.

Next time: the last one. Second dry season ending, the whole cast in the Canopy, and a bill that is larger than it was at the start of all this. The difference is that this time everybody in the room can explain why. Lesson 25 assembles all twenty five lessons, a maturity model, and a ninety day plan for anybody starting from where the jungle was two years ago.

Changing your mind in public, at seniority, is the rarest and most valuable behavior in this entire story. No framework produces it and every framework depends on it.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.