After the meeting in the last lesson, the jungle needed evidence, and discovered it did not have any.
Not because nobody had measured anything. Everybody had measured a great deal. The problem was that all of it had been measured after the capability went live, which meant every comparison was against a version of the old process that existed only in the memories of the animals who had stopped doing it.
It was at this point that the Owl said something that changed the quarter.
“I still have the north territory,” he said.
Everyone looked at him.
What had happened, it emerged, was this. When the rollout was designed, the Owl had produced a phased plan, in the way that Owls do, with the north territory scheduled last for reasons involving a system migration. The migration had been delayed. Nobody had revisited the plan, because nobody revisits a rollout plan once the rollout is going well. So the north territory had spent two full seasons doing the old process, unchanged, while everybody else changed, and their operational data had been collecting the whole time in the ordinary way.
He had, entirely by accident and with no intent whatsoever, been running a control group for eight months.
“Did you know?” asked the Crow.
“I knew they hadn’t gone live,” said the Owl carefully. “I did not know that was interesting.”
Owls have very large eyes. On this one occasion, something had actually been in front of them.
The remembered baseline
Almost every organization begins measuring after deployment. This is not laziness. It is that before deployment nobody yet cares, and by the time somebody cares, the thing they wanted to compare against has been replaced.
So the comparison becomes a comparison against memory, and memory has a consistent bias in this specific situation. People remember the old process as slower and more painful than it was, because what they remember most vividly are the bad instances. Ask a team how long something used to take and you will reliably get a number closer to the worst case than the median.
Which produces benefit cases that are inflated in a direction nobody chose, built by people acting in good faith, and completely undefendable the moment anyone asks where the baseline came from.
The baseline is the whole game. Everything else in this post is technique for when you did not capture one, which is most of the time.
The methods ladder
Four approaches, strongest first. Use the highest rung you can reach.
Randomized holdout. A comparable population that does not get access, chosen at random, for a defined period. This is the cleanest thing available and it is politically the hardest, which is why it is rare. It is also possible for far longer than most organizations assume, because full deployment usually takes longer than anyone plans anyway.
Staggered rollout. Territories go live at different times, and the not-yet-live ones serve as the comparison. This is nearly free, because you are almost certainly already doing it, and the great majority of organizations throw the resulting evidence away by not capturing it. The Owl’s accidental control group was this, discovered eight months late.
Pre and post with controls. Before and after, adjusted for the confounders you can identify. Seasonality, volume changes, staffing changes, process changes shipped at the same time. Weaker than the first two but genuinely usable if you are honest about the adjustments.
Cohort comparison. Heavy users against light users. This is the weakest rung and the most commonly used, because it requires nothing but the data you already have. The problem is that self-selection is doing most of the work. The animals who adopted heavily are not a random sample. They are the ones who were already better at this, or more motivated, or had the kind of work that suited it, and you are measuring that rather than the capability.
The confounders that bite
Four that show up specifically in this domain and that will eat an otherwise sound analysis.
- The process change that shipped alongside. Almost nobody deploys a capability without also simplifying the workflow around it. That simplification often produces more of the benefit than the capability does, and the two are attributed together.
- Novelty. New things get attention and effort. Some of the early improvement is people trying harder because something is new, and it goes away.
- Observation. Teams that know they are being measured perform differently. This is old news and it applies here as much as anywhere.
- Adoption order. Your best people adopt first, reliably. Early results are therefore measured on your strongest performers and will not replicate across the rest.
Making a holdout survivable
The objection to holdouts is always the same and it is not unreasonable. You are denying a useful tool to a group of colleagues in order to generate a chart.
The framing that works, and I have seen it work several times, is to stop describing it as denial. The holdout group is not being deprived. They are the reason the program can prove it works, which is the reason it will still be funded next year, which is the reason everyone eventually gets it. That is true and people respond to it being said out loud.
Three things make it stick. Timebox it, with a date. Guarantee access at the end, in writing. And name the group publicly as the reason the evidence exists, in the same presentation where the results appear. The Mandrill volunteered one of his teams for this in the following quarter, and it did him considerably more good politically than it cost him operationally.
When to measure
Three points, and you need all three.
Week two, which is your peak and which you should record while knowing it is a peak. Month three, which is where the novelty has gone and the real shape starts to appear. Month twelve, which is the number that should drive any permanent funding decision.
The numbers will fall between these points and that is not failure, it is the reference point moving as the improved process becomes normal. What matters is that the organization sees all three, because an organization that only ever sees week two will make every decision on a peak, and an organization that only sees month twelve will conclude the whole thing underdelivered without ever knowing what the early state looked like.
Put the month twelve measurement in the calendar now, while everybody is still enthusiastic enough to agree to it. Nobody schedules that meeting later.
How much rigor is enough
A brief word against overdoing this, since I have just spent a thousand words arguing for evidence.
Match the strength of your evidence to the stakes of the decision. A modest capability serving one team does not need a randomized holdout. Pre and post with a sensible adjustment is fine, and demanding more is a way of making sure nothing gets measured at all because the bar is exhausting.
A decision about whether to deploy something across the organization permanently is a different matter. There the cost of a holdout is trivial against the cost of being wrong, and the fact that it is inconvenient is not an argument.
The failure mode I see most is uniform rigor. Either everything gets the full treatment, which means nothing does, or nothing does, which means everything is anecdote. Tier it.
Three ways this goes wrong
The remembered baseline. Comparing against how long people think it used to take. Ubiquitous, invisible, and fatal the moment anyone asks.
Rollout without capture. Running a staggered deployment, generating a natural experiment for free, and recording none of it. This is the most wasteful thing in this post because the evidence was already being produced.
Measuring once, at the peak. Funding a permanent capability on a week two number, then spending year two explaining why the benefits shrank when nothing shrank except the novelty.
The Field Kit
Concrete things to do this week.
If you sit in the Crow’s chair, ask when the baseline was captured and by whom. If the answer is after go-live, discount the case accordingly and say so out loud, so that the next person to present to you captures one first.
If you sit in the Crocodile’s chair, capture the staggered rollout data. You are already generating it. Recording it costs close to nothing and it is worth more than any dashboard you could build this year.
If you sit in the Mandrill’s chair, volunteer a team as a holdout and get named publicly as the reason the program can prove anything. It is a much better position than it appears from the outside.
For everyone: put the month twelve measurement in the calendar today. It is the only one of the three that nobody will schedule voluntarily and the only one that should drive a permanent decision.
Jungle Lesson 9
You cannot measure a change you did not measure before. Capture the baseline while you are still bored by the idea, because by the time you need it, everyone will remember the old process as being much worse than it was, and you will have no way to prove otherwise.
Next time: the Tortoise looks at everything the Fox has built over the last four lessons and points out, with some satisfaction, that he has accidentally constructed the audit trail she has been asking for since before any of this started. Compliance and the AI lead end up on the same side of the table, which surprises both of them. Lesson 10 closes the second act with the value ledger.
Somewhere in your organization there is probably an Owl sitting on a delayed rollout that nobody has revisited. That is not a project management failure this month. It is the only clean comparison you are going to get.