Scaling AI FinOps | Lesson 21: The Second Year Problem

Somebody found the slide.

It had been in an archive for two years, in a deck that nobody had opened since the quarter it was presented. Somebody was looking for something else entirely and read it out to the room, in the way people do when they find a document that has aged into comedy.

It was titled Conservative Estimate.

What was remarkable about it, once everyone stopped enjoying themselves, was the pattern of what it had got wrong. Every number the Fox had been confident about was wrong, most of them by a wide margin. The volume forecast, the unit cost, the adoption curve, the payback period. All confidently stated, all incorrect.

And the things he had flagged as uncertain, hedged, caveated, marked as needing further work, had turned out to be roughly right. He had been accurate exactly where he had been unsure and wrong exactly where he had been certain.

The Fox took this well, which is the clearest measure of how far he had come. Two years earlier he would have explained it. Instead he asked for the slide to be pinned up in the team area, which it was, and which turned out to be one of the more effective governance interventions of that year.

“The useful part,” he said, “is that I would make exactly the same mistakes again if I did not have this in front of me.”

The second dry season

Everything in the first four acts is about building something. This act is about what happens to it, and the answer is that it decays, in four specific ways that nobody plans for because year one is too busy and year two is too late.

Year two is where AI programs actually die. Not year one, where there is novelty, executive attention, a clean story and permission to be imperfect. Year two has none of those, plus a bill that has grown and a benefits case that has aged.

Why savings evaporate

Configuration entropy. Prompts grow. Retrieval widens. Routing rules acquire exceptions. Every individual change is justified, made by a competent person solving a real problem, and approved by somebody sensible. The aggregate quietly undoes a year of optimization work. Nobody is responsible for the aggregate, which is why it happens.

Workload drift. The distribution of what people ask changes as they learn what the system is good at. Your routing rules were tuned against last year’s traffic mix, which no longer exists. The tuning has not degraded. The thing it was tuned for has moved out from under it.

Success-driven volume. The capability works, so usage grows. Unit economics improving does not stop the total from rising, and the total is what appears in the budget conversation. This is the Lesson 1 inversion arriving again, two years later, at a much larger scale, in front of a more senior audience.

Baseline reset. This is the subtle one. Year two efficiency is measured against year one’s improved baseline. The same absolute gain now looks like a smaller percentage improvement, so it gets deprioritized in favor of things that look more impressive, and the compounding stops.

The savings ratchet

The specific mechanism that catches finance functions, and it mirrors exactly the benefits problem from Lesson 10.

An optimization is delivered. A saving is reported. The saving goes into the cumulative total and is assumed permanent, because that is how savings have always worked in every other category of spend.

Almost none of them are permanent here. The routing improvement erodes as the traffic mix shifts. The prompt trimming erodes as new lines get added after new incidents. The caching gain erodes as somebody restructures a prompt for good reasons and breaks the prefix.

So the correct treatment is identical to the benefit treatment in the value ledger. Every efficiency claim carries a decay assumption and a re-verification date. You will be wrong about the assumption and you will find out on a date rather than in an audit, which is the entire point.

Value drifts too

The other half, and it is the one organizations almost never check.

The benefits case built in year one described a process. Two years later that process has changed, partly because of the capability itself, and partly because processes always change. The comparison that justified the investment may no longer describe anything that exists.

Sometimes this makes the benefit larger. More often it makes it smaller, or simply incoherent, because you are comparing against a way of working that nobody in the building remembers.

Staff turnover accelerates this. The people who did it the old way have moved on. The new people have only ever known the current process, so the improvement is invisible to them, and nobody defends a benefit they have never experienced.

Re-baselining

The discipline that answers all of this, and it is uncomfortable and non-negotiable.

Once a year, re-measure both sides against current reality rather than against the original case. What does a unit cost now. What is it actually worth now, given how the process works now. Then publish the difference between that and what the ledger says.

This is the same reconciliation habit from Lesson 10, applied to the passage of time rather than to a single claim. And it produces the same result: two uncomfortable cycles, followed by an organization that trusts its own numbers more than any of its peers trust theirs.

The alternative is a cumulative savings total that grows every year, describes nothing, and collapses the first time somebody senior asks it a hard question.

Why year two kills programs

Worth stating directly, because it is the thing I would most want a program leader to understand.

Year one is funded by novelty. There is executive interest, a clean narrative, tolerance for imperfection, and an assumption that things will improve. Almost anything survives year one.

Year two has to be funded by evidence. The novelty is gone. The executive who sponsored it has three new priorities. The bill is larger and the benefits case is two years old. And the only thing that will carry the program through is the ability to show what was claimed, what landed, what decayed, and what is being done about it.

An organization that built the ledger and the rhythm in year one walks through this. An organization that did not has no way to defend the program, and it gets cut. Frequently while working, because working and being able to demonstrate that you are working are two different things, and only one of them survives a budget review.

Three ways this goes wrong

The permanent saving. Booked once, never re-verified, quietly gone within three quarters while still appearing in the cumulative total.

Configuration entropy with no owner. Exceptions accumulating, nobody responsible for the aggregate, the estate drifting back to its expensive default state one reasonable decision at a time.

Defending year two with year one’s story. Same slides, aged numbers, a narrative that everyone in the room has heard before and nobody believes twice.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, re-verify last year’s savings before you count them again this year. The number is almost never what the ledger says, and finding that out yourself is considerably better than having it found for you.

If you sit in the Crocodile’s chair, own configuration entropy explicitly. Put a standing review of prompts, retrieval settings and routing exceptions on the calendar. Drift is not an incident, it is a condition, and it needs an operational owner rather than an investigation.

If you sit in the Mandrill’s chair, re-baseline the benefit against how the process actually works now. If nobody in your team remembers the old way, that is your answer about how defensible the original number still is.

For everyone: put a decay assumption and a re-verification date on every efficiency claim, exactly as you would for a benefit. Symmetry between the two sides of the ledger is what makes it trustworthy.

Jungle Lesson 21

Nothing you optimized stays optimized. Prompts grow, exceptions accumulate, traffic shifts, and the baseline you improved becomes the baseline you are judged against. Year one is funded by novelty and year two has to be funded by evidence, which is why the programs that die are usually the ones that were working.

Next time: the last of the Hummingbirds from Lesson 1 finally gets switched off, twenty one lessons later. Nobody can remember what it was for. Its owner left the jungle two seasons ago. It has been running the entire time. Lesson 22 is about decommissioning, and why every enterprise is structurally incapable of it.

If you have a slide from two years ago that you were confident about, go and read it. The pattern of what you got right and what you got wrong is more useful than any forecasting methodology, and it costs you ten minutes and a small amount of dignity.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.