In Practice: AI in the Enterprise | Day 6: The Data Quality Assumption We All Make

There’s a moment in every enterprise AI deployment when someone walks into a meeting with their arms crossed and says: “Our data quality isn’t good enough.”

What they mean is: “I’ve looked at the source systems. Some fields are empty. Some are wrong. Some mean different things depending on which department entered them.”

And then everyone nods solemnly, as if this is a surprising discovery. As if 1980 called and wants its data problems back.

The lie isn’t that you have bad data. Of course you do. The lie is the myth of good enough.

Most organizations operate under a belief that’s almost never stated explicitly, but it sits there like furniture in the room: if you can get your data to 90% quality, you can build a functional AI system. Maybe 85%. Definitely at 95%.

This is false.

It’s not false because data quality doesn’t matter. It’s false because you can build a functional AI system with pretty broken data—you just can’t build one you can trust to make decisions at scale.

I’ve watched teams spend eight months fixing data pipelines. Getting the nulls down. Standardizing the schemas. Running reconciliation checks. And when they get the system into production, it performs fine on metrics. It produces outputs. But when a human looks at what it’s actually recommending—especially in edge cases, in the long tail—there’s something off about it. Not catastrophically, not always. Just… off.

The system learned the pattern, including the pattern of your data quality problem.

Here’s what’s actually happening underneath. Your data has three types of defects:

Structural defects are easy. Missing values, wrong data types, things that break schemas. Your tools catch these. You fix them. Life goes on.

Semantic defects are harder. A field labeled “revenue” that means net revenue in one business unit and gross in another. A “customer type” that was coded one way three years ago and a different way now, but nobody updated the historical data. A date field that’s sometimes a transaction date and sometimes a creation date, depending on which system it came from. These are invisible to structural checks. They look clean on a data quality scorecard. And they teach your model nonsense.

Temporal defects are the ones that kill you. Your historical data is, by definition, old. It was collected under different business conditions, different processes, different market dynamics. An AI system that learns from it is learning from a past that no longer exists. You’ll discover this about six weeks after your system goes live, when its predictions start diverging from reality.

The real question isn’t: is our data good enough?

It’s: what decisions are we letting this system make, and what defects can we actually live with?

This is where most conversations break down. Because it forces you to be specific. Specific means uncomfortable.

A recommendation engine that occasionally suggests a slightly irrelevant product? You can live with that. Your customers see a mediocre recommendation, they ignore it. Cost of that failure: low.

A system that scores credit risk or flagged fraud, and it’s learned to be structurally inconsistent about what “fraud” means? You can’t live with that. You’ll catch it in compliance review three months in and have to rebuild it.

The practical move isn’t to announce you need “better data quality.” That’s governance theatre.

The practical move is to map your data defects against the specific decisions you’re making. Ask yourself: where does this data come from? Who was incentivized to enter it, and what were they incentivized to do? What changed three years ago in the system that collected this? Where is it inconsistent, and what does that inconsistency mean for what we’re trying to predict?

Then: deliberately reduce the scope of decisions you let the system make until the defects stop mattering.

You might keep the system from making binary decisions and make it advisory instead. You might restrict it to certain business units until you understand why the data behaves differently there. You might run it in shadow mode longer than you planned because the temporal defects are real and you want to see how it drifts.

This is boring. It’s not a feature announcement. It won’t make the quarterly exec summary.

But it’s the difference between a system that hums along in production for three years and one that breaks down in month eight when someone realizes it’s been quietly making recommendations that contradict its own logic because your semantic defects taught it two contradictory patterns.

Your data quality problem isn’t that you need better data. It’s that you need specific knowledge about where bad data actually matters, and you need the organization discipline to say: not yet. Not there. Not until we understand this.

That’s harder than buying a data quality tool. And it’s exactly the kind of hard that distinguishes deployments that compound value over time from the ones that become cautionary tales.

In Practice: AI in the Enterprise | Day 5: A CFO, a VP of AI, and a Compliance Officer Walk Into a Decision… Who Actually Decides?

This isn’t a joke. It’s how many AI deployment decisions actually happen in enterprises, and it ends the same way: with nobody being entirely clear about who made the call.

The CFO’s concern is straightforward: if this fails, how much does it cost us? They want to understand the financial consequences. They’re thinking about downside risk, capital allocation, shareholder communication if something goes wrong.

The VP of AI is thinking about capability and execution. Can we actually build this? Will it work? What will we learn? They’re thinking about staying competitive, building the team’s skills, proving value. They want to move.

The Compliance Officer is thinking about exposure. What if a regulator asks about this? What if a customer complains? Can we defend this decision? They’re thinking about policies, documentation, the audit trail. They want to be careful.

These aren’t contradictory perspectives. But they’re different perspectives, and they weight risks differently. The CFO is most concerned about economic risk. The VP is most concerned about execution risk. The Compliance Officer is most concerned about regulatory risk. When you put them in a room to decide whether to deploy an AI system, someone has to make the call. And often, nobody actually does.

What usually happens instead is consensus seeking that masks disagreement. You leave the meeting with implicit agreement that “we’re moving forward,” but the CFO is thinking “we’re moving forward if the economic model works out,” the VP is thinking “we’re moving forward and we’ll figure out the details,” and the Compliance Officer is thinking “we’re moving forward and we’ll handle the risks.” These statements sound like agreement. They’re not. They’re three different decisions that will conflict the moment execution requires actual tradeoffs.

Then something happens—the model drifts, a customer complains, a regulator asks a question—and suddenly it’s clear that you never actually resolved the disagreement. You just deferred it.

The Three Types of Risk These People Actually Represent

To understand why this happens, start with what each person is actually accountable for.

The CFO is accountable for financial performance. If the AI system costs more than it saves, they answer to the board. If it creates unexpected costs, they’re explaining variance to investors. Economic risk is their domain.

The VP of AI is accountable for delivering on the AI roadmap. Building capabilities, shipping systems, enabling the business. Execution risk is their domain. They’re being evaluated on delivery and impact, not on managing risk.

The Compliance Officer is accountable for regulatory compliance and risk management. If something goes wrong and there’s no documented governance, they’re explaining to auditors and regulators. Governance and procedural risk is their domain.

These aren’t competing priorities. They’re different dimensions of risk. The problem is that your governance structure probably treats them as competing. It probably puts them in the same meeting, assumes they’ll reach consensus, and then acts surprised when they don’t.

The real issue is that nobody is explicitly accountable for integrating these perspectives and making the actual decision. The CFO can’t decide alone because they don’t understand execution risk. The VP can’t decide alone because they’re not responsible for financial outcomes. The Compliance Officer can’t decide alone because they’re not responsible for business outcomes. So you have three people none of whom is empowered to make the call.

This is a different governance problem than the ones in previous pieces. In Days 1-2, I was focused on the structural problem: nobody owns the decision. Here we’re looking at the substantive problem: when you have multiple valid perspectives and they don’t align, how do you actually decide?

What Happens When You Pretend Consensus Exists

Most organizations try to solve this through consensus. The theory is: if everyone agrees it’s a good decision, it’s a good decision. If people disagree, talk until they agree.

This works for some decisions. It doesn’t work for decisions where people have legitimately different risk appetites.

Say the AI system will cost $500K to build and maintain, and it will save $1.5M annually if it performs as modeled. The CFO looks at this and says: “If performance is as modeled, great. If it underperforms by 20%, we still have positive ROI, so I can support this.” That’s a CFO’s way of agreeing.

The VP of AI looks at the same numbers and says: “I can build this, and I believe the model will work well.” That’s a VP’s way of committing to execution.

The Compliance Officer looks at the same numbers and says: “I need to see the governance around this, but the financial case looks sound.” That’s a Compliance Officer’s way of saying they’ll proceed carefully.

All three have “agreed.” But the CFO agreed subject to financial performance. The VP agreed subject to being able to execute. The Compliance Officer agreed subject to governance being in place. If any of those conditions isn’t met, the agreement falls apart. But nobody documented the conditions. So everyone walks out of the meeting thinking they’ve committed to moving forward, when what they’ve actually committed to is moving forward if certain things happen.

Then execution reveals that governance isn’t as clear as the Compliance Officer wanted, or execution is harder than the VP expected, or the financial model is too optimistic. Now you have a decision you thought was made, coming apart under the stress of reality.

The Actual Decision-Making Process That Works

The organizations I’ve seen navigate this successfully have a structure that acknowledges this complexity rather than hiding from it.

They explicitly assign decision authority to one person. Often this is a business leader—a VP of Product, a business unit leader, a Chief Operating Officer. Someone who’s accountable for outcomes, not just for process. That person is explicitly accountable for the deployment decision.

They structure the input so that person gets clear perspective from all angles. Not consensus—perspective. The CFO provides financial analysis and financial risk tolerance. “Here’s what happens if the model underperforms by 10%, 20%, 30%. Here’s the financial exposure.” The VP of AI provides execution assessment. “Here’s our confidence in our ability to build this. Here’s what could go wrong technically. Here’s how we’d mitigate it.” The Compliance Officer provides governance and regulatory assessment. “Here’s what we need to document. Here’s what regulators might ask. Here’s what we can defend.”

The decision-maker’s job is to integrate these inputs, understand the tradeoffs, and make a call. Not “do we all agree?” but “given what we know, should we move forward?”

They document the decision explicitly. Not “we decided to deploy,” but “we decided to deploy this system, accepting these financial risks, with this mitigation plan, having documented this governance, knowing that these scenarios could happen and here’s how we’ll handle them.” That decision document becomes the reference point when things change. If something goes wrong, you can point to what you understood and accepted at the time.

They assign someone to own ongoing accountability. Usually the same person who made the deployment decision. That person is accountable not just for the decision being made, but for it being monitored and for adjustments when things change. They’re the person who can actually pull the system if something goes wrong.

The Specific Governance Moments Where This Breaks

In my experience, this breaks down in specific places:

The decision meeting that doesn’t resolve anything. You’re in a room. Three perspectives. No explicit decision authority. You talk until everyone is tired, and then you say “okay, sounds like we’re aligned on moving forward.” You’re not aligned. You’re just tired. Write down what you’ve agreed to. Document the conditions and assumptions. Have the decision-maker say explicitly: “Here’s what I’m deciding, and here’s what I’m accepting by making that decision.”

The governance framework that exists but doesn’t actually gate anything. You have a checklist. Model validation: check. Fairness assessment: check. Compliance review: check. And now you’re automatically approved to deploy. This is governance theater. Someone has to actually make a decision based on the checklist. Not “has the checklist been completed,” but “given what the checklist shows, should we deploy?” If the answer is always “yes,” you don’t have governance. You have a process.

The risk discussion that’s all upside and no downside. You’re talking about how much value this will create. Nobody’s talking about what happens if it doesn’t. The CFO should be laying out financial downside. The VP should be laying out execution risk. The Compliance Officer should be laying out governance and regulatory risk. If everyone is only talking about upside, you haven’t actually thought about risk. You’ve just decided you like the idea.

The monitoring plan that nobody actually monitors. You’re going to “track accuracy” and “monitor for bias” and you’ve assigned it to a team. That team has other priorities. Months later, nobody’s looking at the metrics. And the first time something goes wrong is the first time anyone actually paid attention. Assign actual, prioritized monitoring to someone. Not “the data science team will keep an eye on it,” but “person X is responsible for checking this metric weekly, and person Y is responsible for deciding what to do if the metric moves.”

The escalation that doesn’t actually escalate. Something unexpected happens with the model. The monitoring team sees it. They send an email. It gets lost. Or it goes to a stakeholder who doesn’t have the authority to make decisions. Or it gets escalated to a committee that meets quarterly. By then it’s late. Build an escalation path where an observation actually reaches someone with authority to act, and with enough urgency to matter.

How This Looks in Practice

Here’s a real-world example of what I mean (composite, not a specific case):

You’re deploying an AI system for credit decisions. The CFO looks at the model and says: “If this achieves 85% approval rate with a default rate of 4%, I can fund this. If default rate goes above 5%, we need to recalibrate or stop.” That’s the financial decision rule.

The VP of Risk (who owns credit) says: “I can defend these decisions if I can explain them. If we’re using a black-box model and regulators ask why we denied someone credit, I need an explanation. I need model documentation and decision-level traceability.” That’s the governance decision rule.

The Compliance Officer says: “Fair Lending is a regulatory focus. We need to demonstrate that our approval rate is not substantially different across demographic groups, or if it is different, we need to explain why based on credit risk factors, not demographic factors.” That’s the regulatory decision rule.

The decision-maker’s job is to say: “We’re deploying this system. The target is 85% approval with 4% default. We’re committing to documenting decisions and having explainability. We’re targeting demographic parity adjusted for credit risk. If either of the first two gets violated, we recalibrate. If the third gets violated, we stop. Person A monitors the approval rate and default rate weekly; person B monitors the decision documentation and explainability; person C monitors the demographic parity. If any metric moves outside acceptable range, person A/B/C escalates to me for decision.”

That’s a decision. Not a consensus that hides disagreement. A decision that integrates different perspectives and lays out what happens next.

The Board Question

When the board asks “why did you deploy that AI system?” the answer should be specific:

“We evaluated the deployment against financial risk, execution risk, and governance risk. The financial model showed positive ROI with acceptable downside. We had sufficient technical capability to execute. We had governance and regulatory framework in place to defend the decision. We assigned clear ownership and escalation. We made the decision deliberately, documenting what we understood and accepted.”

That’s what a governance decision sounds like. Not “everyone agreed,” not “the model was validated,” not “it was on the roadmap.” But “we thought through the decision, we integrated perspectives from across the organization, and we were willing to own the consequences.”

When you have that clarity—about who decides, what they’re deciding, what they’re accepting, what happens if things change—the CFO, the VP, and the Compliance Officer stop needing to reach consensus on something they never fully agreed about in the first place. They each provide clear perspective on their domain, someone integrated that perspective into a decision, and they all know what was decided and why.

That’s how you actually handle multi-stakeholder AI deployment decisions. Not perfectly. But deliberately.

In Practice: AI in the Enterprise | Day 4: Why Your Model Validation Process Is Probably Measuring the Wrong Thing

Most organizations answer this question with metrics. Accuracy. F1 score. AUC. Precision and recall. So they build validation processes optimized for those numbers. They backtest them. They test on holdout datasets. They measure them on new data. And then they watch the metrics carefully after deployment.

Here’s the problem: you’re measuring whether the model works. You should be measuring whether the decision it produces is acceptable.

These are different things.

A model can be technically accurate and still produce decisions that violate your risk tolerance. You can have high F1 scores and low decision quality. You can reduce bias metrics to perfect parity and still fail because you’re optimizing for the wrong fairness measure, or you’re achieving fairness in a way that destroys business value, or you’re failing on a different fairness measure you didn’t think to check.

The organizations that actually manage model risk well aren’t the ones with the most sophisticated validation metrics. They’re the ones that ask a different question first: what decision quality do we require? And then: how do we know if we’re achieving it?

The Accuracy Trap

Start with accuracy, because it’s the easiest one to get wrong.

You’ve trained a model to predict something. It’s 95% accurate on test data, 94% on a new holdout set, 93% after the first month in production. These are good numbers. And they’re technically meaningless for understanding your risk.

What matters is what happens when the model is wrong. If the model is predicting customer churn, and it’s 93% accurate, that means it’s wrong 7% of the time. That 7% breaks into two categories: people it predicted would churn who didn’t (false positives), and people it predicted wouldn’t churn who did (false negatives). These aren’t symmetric costs.

A false positive means you treat a loyal customer as a churn risk. You might offer them discounts, special treatment, outreach. That’s usually not expensive. It’s annoying marketing waste.

A false negative means you miss a customer who’s actually going to churn. You don’t intervene. They leave. That’s actual business loss.

So from a business perspective, 93% accuracy in a churn model might be a disaster (if the 7% errors are mostly false negatives) or acceptable (if they’re mostly false positives).

Your validation metrics don’t tell you this. You measure accuracy, and it’s high. You’ve validated the model successfully. And then in production, you’re failing at the business problem you were trying to solve, because you optimized for the wrong metric.

This scales up. In lending, false negatives (denying credit to someone who would have repaid) cost you revenue. False positives (approving credit for someone who defaults) cost you loss. The acceptable balance depends on your risk appetite, your business model, your regulatory environment, your customer expectations. Not on what the model’s F1 score is.

In hiring, false positives (rejecting candidates you would have wanted) destroy your talent pipeline and waste recruiting costs. False negatives (hiring people who don’t succeed) destroy your teams. The acceptable ratio depends on your hiring philosophy, your training pipeline, your attrition rates, your regulatory obligations. Again: not on the accuracy metric.

Most validation processes don’t actually think about this. They measure accuracy as a proxy for quality, and they assume that if the metric is good, the decision is good. It often isn’t.

The Fairness Metric Confusion

This gets worse with fairness, because fairness metrics sound more sophisticated but are actually more fragile.

You’ve committed to making fair decisions. Great. Now: what does fairness mean?

Equal accuracy? The model predicts defaults equally accurately for applicants of all races, genders, ages. Sounds good.

Except: if default rates are actually different across groups (maybe due to historical discrimination, maybe due to real difference in risk, depends on the domain), then equal accuracy means equal Type I and Type II error rates, which means unequal impact. You approve proportionally more loans to the group with lower baseline default rate.

Equal impact? The model rejects the same percentage of applicants from each demographic group. Sounds fair.

Except: if rejection rates should actually be different (if the groups have different risk profiles), then equal impact means you’re being statistically inaccurate, and you’re probably violating fairness in some other direction. You’re approving loans you shouldn’t approve, or rejecting people you should approve.

Equal opportunity? The model is equally likely to identify a non-defaulter as non-defaulter, regardless of demographic group. This is sometimes called “equalized true positive rate.”

Except: this assumes that not defaulting is equally distributed across groups, which it might not be. And it doesn’t address false positive rates. And it’s a different fairness metric than the last two.

There are eight major fairness frameworks, and they’re mutually incompatible. You can optimize for one and fail on all the others. And there’s no mathematical way to optimize for all of them simultaneously.

Most validation processes pick one metric, measure it carefully, optimize for it, and call that “ensuring fairness.” Then they’re surprised when their model is challenged on a different fairness dimension. Or when their fairness metric is high but the decisions are still systematically wrong for a particular group.

The problem is the same as with accuracy: you’re measuring a proxy, and assuming the proxy is the outcome you care about. It usually isn’t.

What Actually Matters: Decision Quality

Here’s what validation should actually be checking: does this model produce decisions I’m willing to make and defend?

That’s a different question. It requires you to think through what you’re actually trying to accomplish with this model, what the constraints are, and what outcomes you need to be able to defend.

For a churn model: can you identify which customers are most likely to leave, well enough that your intervention strategy is cost-effective? Can you explain why you’re intervening on customer X but not customer Y? Can you explain what happens if you’re wrong? Can you accept the distribution of errors (false positives and false negatives) that you’re creating?

For a lending model: can you identify creditworthy applicants reliably enough that your default rate is acceptable? Can you defend the decisions you’re making about particular applicants? If the model says “reject” but you override it and approve them, how often do they succeed? If the model says “approve” but you reject them anyway, how often would they have succeeded? Can you accept the financial consequences of the error rate?

For a hiring model: can you identify candidates who will succeed in the role? Can you explain why you screened out a particular candidate? If the model says someone won’t succeed but they do (in your competitors’ organizations), how do you reconcile that? Can you accept the representation outcomes you’re creating?

These questions require business judgment and domain knowledge, not metrics. You can’t reduce them to a single number. But they’re the actual validation question.

Building the Right Validation Process

If you want to actually validate whether a model is acceptable to deploy, you need a different process:

First: Define decision quality requirements. Not metrics—requirements. What outcomes do we need this model to produce? What are we willing to accept? What are we not willing to accept? For a lending model: “We need to identify 80% of creditworthy applicants without exceeding a 3% default rate.” For a churn model: “We need to identify 60% of likely churners; we’re willing to spend on interventions that have a 50% success rate.” For a hiring model: “We need candidates with >70% 2-year success rate; we want at least 30% representation of underrepresented groups in our hire cohort.” These are business requirements, not technical metrics.

Second: Measure whether you’re meeting them. This is where metrics come in, but they’re metrics aligned with your actual requirements. Not “what’s our F1 score,” but “of the people we identified as likely to churn, did 50% actually churn?” Not “what’s our fairness metric,” but “do we have the representation we said we wanted? And at what cost to accuracy?” You’re validating that you’re meeting the specific requirements you set, not optimizing for proxies.

Third: Understand the tradeoffs you’re making. Every model involves tradeoffs. Better accuracy at the cost of fairness. Higher profit at the cost of customer risk. Faster decisions at the cost of accuracy. Most organizations don’t deliberately articulate these tradeoffs. They just happen. Build validation that makes them explicit: “If we deploy this model, we gain X in business value and we accept Y in risk. Here’s what Y looks like in concrete terms.” Make sure the person approving the deployment actually understands what they’re approving.

Fourth: Establish the failure threshold—before deployment. What would cause you to pull this model? Not “we’ll know it when we see it,” but “if [metric] exceeds [threshold], we reconfigure the model or take it offline.” This isn’t about identifying all problems; it’s about identifying the problems that matter most. You might not care if your churn model’s accuracy drifts from 93% to 90%. You probably do care if it starts making decisions that violate your fairness requirements. You definitely care if the cost per intervention increases beyond what makes the program economically viable.

What Happens When Validation Aligns with Decision-Making

Organizations that actually validate well have one thing in common: they’ve spent time thinking through what decision quality looks like for their use case, before they build the model. The validation process then checks whether they’re achieving it.

This sounds obvious. It’s not what most organizations do.

Most build the model first, validate against metrics, and then try to justify deployment based on the metrics. They work backward from the numbers they have to an acceptable story about what those numbers mean. Sometimes the story holds up. Often it doesn’t.

When someone says “our validation shows this model is ready,” what they usually mean is “we have metrics that look good.” They usually don’t mean “we have confirmed this model produces decision quality we’re willing to defend.” These are different things.

The reason this matters for governance and risk is straightforward: when something goes wrong with the model, someone will ask you to justify the deployment decision. “This model is making this decision. Why did you think that was acceptable?” Your answer can’t be “our F1 score was 0.87.” Your answer has to be “we required [X outcome], we validated that we achieved it, we understood the consequences, and we decided it was acceptable.” That’s an answer you can defend. The other one isn’t.

Your validation process isn’t just a technical gate. It’s the evidence for your governance decision. Make sure it’s measuring what actually matters.

In Practice: AI in the Enterprise | Day 3: The Compliance Gap Nobody’s Talking About: What Regulators Actually Expect vs. What You’re Preparing For

Myth: Regulators are coming after your AI with specific technical requirements.

Reality: They’re coming after you for governance failures that AI just happens to amplify.

This distinction explains why so many enterprises are spending heavily on the wrong things.

Over the past year, I’ve talked to regulatory and compliance leaders at a dozen organizations. The conversation is remarkably consistent. They’re building AI governance programs that look sophisticated: model risk frameworks, fairness audits, documentation requirements. But when they actually talk to regulators, whether in an exam, an inquiry, or a pre-filing consultation, the questions are almost never about the AI itself.

The regulators ask:

  • How do you know what your model is doing?
  • Who decided this was acceptable to deploy?
  • What happens when something goes wrong?
  • How do you know it’s degraded if you’re not looking?
  • Did you consider this particular risk?
  • Can you actually explain the decision it made in that case?

These aren’t technical AI questions. They’re governance questions. They’re accountability questions. They sound like questions a board would ask about any high-risk business process. Because that’s what regulators think AI is: a high-risk business process, not a technical marvel.

Most organizations have gotten this backward. They’re building technical AI governance, fairness metrics, model documentation, validation frameworks—as though the regulatory risk is technical. It isn’t. The regulatory risk is organizational. Did you think about it? Did you decide it was acceptable? Can you prove it?

What Regulators Actually Care About

Start with what they don’t care about (or at least, not what you think).

They don’t care whether you use logistic regression or a transformer model. They don’t care about your accuracy metrics. They don’t care whether you’ve achieved parity in your fairness score. They don’t even care, technically, whether your model is biased, they care whether you knew it might be biased and whether you decided that was acceptable.

This is crucial. Regulatory failure isn’t “we had a biased model.” It’s “we had a biased model and either we didn’t know it could happen, or we knew it could happen and deployed it anyway without disclosing it or managing it.”

Regulators work from an assumption: you have obligations to your customers, shareholders, or society. You make a decision to deploy AI. If that decision creates a risk of unfair treatment, of customer harm, of financial loss, of regulatory violation, you’re accountable for that decision. Not for the model being perfect, but for the decision being made with appropriate information and oversight.

This is why compliance teams often feel like AI governance is a moving target. They’re being asked to regulate decision-making, not technology. And you can’t write a regulation that says “make good decisions about AI” the way you can write a regulation that says “maintain a capital ratio of X.” So regulators fall back on the question: how do you know you’re making good decisions? Document your process. Show your thinking. Prove you considered the risks.

The compliance framework regulators actually want to see is the framework you’d use for any major business decision that carries significant risk.

Where Enterprises Overinvest (The Wrong Things)

I’ve watched organizations spend significant resources on governance activities that, if challenged by a regulator, would do almost nothing to defend their decision-making.

Building fairness metrics without any process for deciding whether the level of fairness is acceptable. You measure bias, but who decided “this level of disparity is okay, and here’s why”? Nobody. You just built a dashboard.

Creating model documentation templates that describe what the model does, but not why the business decided it was worth doing. A regulator looks at your documentation and asks “okay, but who approved deploying this knowing it could make this particular decision incorrectly 3% of the time?” You don’t have an answer because you never framed it as a decision.

Implementing validation frameworks that test whether the model works, but not whether it’s solving the right problem. You can prove the model is accurate. You can’t prove the business thought through the consequences.

Putting fairness audits in place after deployment and calling it governance. Audits are useful, but audits are reaction. They tell you what went wrong. They don’t tell regulators you thought it through before you took the risk.

These activities aren’t wrong. But they’re supporting infrastructure, not the core governance question. The core governance question is: did you make this decision deliberately, with appropriate oversight, knowing the risks, and documenting your reasoning?

What Regulators Actually Expect to See

When a regulator examines an AI program, the framework they’re using is decades old. It’s the same framework they use for credit decisions, lending practices, capital allocation. Does the organization have:

  • Clear governance structure. Is there someone responsible for AI decisions? Not someone overseeing audits—someone who actually approved the deployment and is accountable for it. They’ll trace reporting relationships. They’ll ask how the decision got made.
  • Risk assessment before deployment. Did you think about what could go wrong? Not “was the model tested” (that’s validation). But “did you consider whether this model could produce unfair outcomes, or incorrect decisions, or customer harm, and did you decide the benefit was worth the risk?” They’ll look for documentation of this thinking.
  • Clear criteria for continuing operation. You’ve deployed the system. How do you know if it’s still acceptable? What thresholds, metrics, or observations would cause you to pull it? They’re asking: did you think about failure scenarios before you had to deal with them? Or are you just running it until someone complains?
  • Evidence of oversight and control. Is someone actually monitoring this, or did you set up an automated system and forget about it? They’re looking for evidence that a human being with responsibility is actually paying attention.
  • Ability to explain decisions. Pick a customer outcome. Tell us why that happened. Can you trace the decision through your process? Can you explain the model’s logic? Can you show that this outcome is consistent with how you use the model overall? This matters even if (especially if) the model is black box.

This is not a technical framework. This is an organizational framework. It applies to loan decisions, hiring decisions, insurance underwriting, investment decisions, and now AI decisions. The fact that a machine is involved in the decision doesn’t change the fundamental regulatory expectation: you thought about it, you decided it was acceptable, you documented that thinking, and you’re overseeing it.

The Specific Compliance Gaps

Here’s where most organizations are vulnerable:

  • Decision documentation. You can explain what the model does. You probably can’t explain why someone in your organization decided it was acceptable to deploy a model that makes this decision this way. Write that down. For every major AI system, there should be a document that says: “We considered deploying this model. It will make decisions about [X]. We know it could [create these risks]. We decided to deploy it because [business rationale]. We are accepting [specific consequences]. We will monitor for [these indicators]. If we see [these outcomes], we will [take these actions].” Most organizations don’t have this. Most regulators would expect it.
  • Ownership and delegation. The same person accountable for the decision should be accountable for ongoing oversight. If you delegate monitoring to a team, you’ve still got the accountability at the top. But that line has to be clear and documented. Regulators will ask “who is accountable for this system?” If the answer is a data science team, you’ve revealed that you don’t have governance. The team is technically supporting a system, but nobody with P&L accountability is responsible for it.
  • Consideration of alternatives and harm. You deployed this AI system. Did you consider not deploying it? Did you consider less risky alternatives? You don’t have to choose the lowest-risk option, but you have to show you considered it. This is especially critical if there’s an alternative (a simpler model, a human review, a rules-based system) that would have lower regulatory risk. You don’t have to choose it, but you have to show you knew about it.
  • Threshold and escalation. When do you pull this system? Not theoretically—actually, what’s the threshold? “If accuracy drops below X” or “if we detect [pattern]” or “if we receive more than Y complaints about this.” Regulators know you won’t have a perfect answer, but they want to know you thought about what failure looks like rather than just discovering it after customers are harmed.

What This Means for Compliance Teams

If you’re building an AI governance program, prioritize the governance questions first. Not because fairness metrics or model documentation don’t matter, but because they’re supporting infrastructure. The core governance question is: did the organization make this decision deliberately?

That means:

  • Work with the decision-maker (whoever that is in your organization) to document the reasoning behind each major AI deployment before it happens, not after. What are we trying to accomplish? What’s the risk? Who decided this was worth it? This is accountability documentation, not technical documentation.
  • Make sure oversight is actually assigned to someone with real authority. Not a committee reviewing outcomes, but someone who can actually make changes. Regulators will ask about the escalation path. If everything escalates to a committee that meets quarterly, you’ve just shown that you don’t actually have real oversight.
  • Define what failure looks like. For each major system, identify: what indicators would suggest this system is no longer operating acceptably? Not “what data would we want to look at,” but “what specific observations would cause us to take it offline or reconfigure it?” And assign someone to actually look at those indicators regularly.
  • Be transparent with regulators about what you don’t know and how you’re managing it. “We’ve deployed an AI system and we’re monitoring X, Y, and Z” is strong. “We’ve deployed an AI system and we’re monitoring what we can” is weak. If you say “we don’t fully understand how this system behaves in edge cases,” but you’ve mitigated it with human review or conservative thresholds, that’s defensible. If you’ve deployed it and you don’t understand what it does and nobody’s overseeing it, that’s not defensible.

The Regulatory Trend

Regulatory frameworks for AI are still emerging, which creates urgency. When frameworks crystallize, organizations that have been building decision accountability from the start will have a much easier transition than organizations that have been building technical governance and hoping it satisfies the regulatory need.

The organizations building fairness audits and model documentation aren’t wrong. But they’re building infrastructure for a governance system they haven’t actually built yet. The core system—clear decision authority, deliberate deployment decisions, documented reasoning, real oversight, has to come first.

When a regulator asks about your AI program, they’re going to ask about governance first, and technology second. If you can’t explain who decided this was acceptable and why, all your fairness metrics and validation frameworks are supporting a house with no foundation.

That’s the compliance gap nobody’s talking about. You’re preparing for technical audits. Regulators are looking for organizational governance. These are different things.

In Practice: AI in the Enterprise | Day 2: When the Board Asks “Who Owns This AI?”, Your Answer Reveals Everything

Your org chart is a lie about how AI actually gets made. Not intentionally. But fundamentally.

Pull out the structure your company publishes. Find the AI or data science team. Now trace the lines. Probably reports to VP of Engineering. Or CTO. Or maybe there’s a Chief Data Officer. Clean lines. Clear hierarchy.

Now describe what actually happens when someone builds an AI system in your organization.

The data science team trains the model. But they didn’t pick the business metric it optimizes for, someone in product did, or maybe operations. They didn’t decide whether to deploy it, that’s usually a product decision, sometimes influenced by finance. They’re not monitoring what it does in production, operations handles that, or risk management, or sometimes nobody specifically owns it. They certainly aren’t making the call if something goes wrong.

So when the board asks “who owns this AI?” and you point to the org chart, you’re not lying. You’re just describing something that has nothing to do with how the decision actually gets made.

This gap, between formal ownership on paper and actual accountability in practice, is where most AI governance breaks down.

The Difference Between Ownership and Accountability

Start with a definition. “Who owns it” in a traditional sense means: whose budget, whose performance evaluation, whose responsibility if it fails. On an org chart, that’s usually clear. But ownership in the org chart doesn’t mean accountability for outcomes.

Accountability means: you made the decision to deploy this, you understood the risks, and if something goes wrong, you face the consequences. You’re the person who can’t pass the problem to someone else.

In most organizations, AI systems have plenty of ownership but almost no accountability. The data science VP owns the team. But the product leader who defined what “good” meant owns the business decision. The ops leader who deployed it owns the infrastructure. The compliance team owns the audit trail. Finance owns the budget impact if it goes wrong.

Everyone owns a piece. Nobody owns the whole thing.

This creates a specific organizational pathology: distributed blame. Something goes wrong—a model makes a bad prediction, or shows bias, or violates a regulatory expectation. Now it’s “well, the data was the problem” (data engineering’s issue), or “the business case was wrong” (product’s issue), or “we didn’t monitor properly” (operations’ issue). Everyone points at a piece of the system they don’t own, and nobody has to fully own the failure.

Compare this to how a traditional business decision gets made. If a CFO approves a capital expenditure and it goes bad, they own it. They were accountable. They had authority to say no. They understood what could go wrong. That structure, clear authority, clear accountability, clear consequence, is what actually drives good decision-making.

AI hasn’t broken accountability as a concept. It’s just exposed that most organizations never gave anyone accountability for AI decisions in the first place. They distributed ownership across multiple teams and called it governance.

Why Org Charts Fail for AI

The problem starts with how data science evolved in most enterprises.

Early data science was an analytics function. It reported to someone, usually an engineer, or a data leader under finance or analytics. Small teams. Answering specific questions. Then AI became strategic, and suddenly these teams were making decisions that could reshape customer experiences, cost millions of dollars, or create regulatory exposure. But the organizational position didn’t change. The team still reported to the same place. Still had the same scope on paper.

Except now they were making much bigger decisions, across multiple business units, with risks that weren’t on their radar.

Meanwhile, product, operations, and compliance weren’t designed to own AI decisions either. Product teams decide what features to build, not how to train models. Operations teams run infrastructure, not model performance. Compliance teams audit outcomes, not design systems before they’re deployed. So you end up with a situation where the people technically closest to the decision (data science teams) don’t have the authority or perspective to own it, and the people who have the authority (business unit leaders) aren’t positioned to understand the technical implications.

This is why “building an AI governance committee” doesn’t actually solve it. Committees are great for coordination. They’re terrible for accountability. Everyone on the committee has other priorities and other people they’re accountable to. Decisions get diffused. Accountability dissolves.

The org chart assumes that ownership flows upward, one person in the hierarchy is ultimately responsible. AI breaks that assumption by making decisions actually horizontal. You need product perspective, technical capability, risk understanding, and compliance knowledge simultaneously. No single person has all of that.

What Accountability Actually Requires

If you want clear accountability for AI decisions, you need three things on paper that match three things in practice.

First: A clear decision-maker. Not a committee. One person. This person needs to have the authority to say no to deployment. The authority to pull a system that’s not working. The authority to redirect resources. That person probably needs to be a business leader, the person whose P&L or customer outcomes are affected, not a technical person. Because the decision being made isn’t “is this technically possible,” it’s “should we take this risk given our business situation.”

Second: Clear information at decision time. That decision-maker needs to know what the model does, what it can go wrong, what the business case is, what the regulatory exposure is, and what monitoring will tell us if it’s degrading. They need this before they approve. Not after. Not in a report six months later. At decision time. That means the technical teams, product teams, and risk teams have to feed information to that decision-maker in a structured way. Not “here’s what we found,” but “here’s what you need to decide.”

Third: Real authority over what happens after. This is where most organizations fail. The decision-maker approves deployment. Then operations runs it. Risk monitors it. Finance tracks costs. Product owns the feature. Nobody has the authority to actually change what happens based on what’s learned. You need that decision-maker, or someone they explicitly delegate to, to have the authority to reconfigure, halt, or redirect based on post-deployment information. Otherwise, the decision-making authority is theatrical.

The Signals You’re Getting It Wrong

Here are the patterns that show up when accountability is actually distributed rather than clear:

  • You have a data science team reporting to engineering and a separate AI ethics team reporting to compliance, and they don’t actually coordinate before deployment because they’re in different chains of command.
  • You have a governance committee that reviews AI projects, but the committee has no power to block deployment, it exists to document that review happened.
  • You’re using three different job titles for people doing similar roles (because the org structure doesn’t actually match the work).
  • When something goes wrong with an AI system, the first five conversations are about whose fault it was, not about how to prevent it next time.
  • You have a Chief Data Officer and a VP of Product making different decisions about the same system because they’re optimizing for different metrics.
  • Your risk and compliance teams are most engaged after deployment, not before it.

What Actually Works

The organizations I’ve seen successfully clarify this usually make three changes:

  • They explicitly assign decision authority to a single person, usually someone in the business (product, operations, or line of business) with real budget and outcome accountability. That person owns the deployment decision.
  • They structure information flow so that person gets input from technical, product, risk, and compliance teams before the decision, not after. This usually means regular design reviews with clear documentation of what each team validated.
  • They give that decision-maker or their delegate ongoing authority to act if something changes. No separate approval needed to reconfigure or halt. The authority flows from the decision-maker, not upward through hierarchy.

This doesn’t require changing your org chart. It just requires being explicit about who actually decides and making sure the org chart doesn’t contradict that. When the board asks “who owns this AI?” the answer should be one name. Not a team. Not a committee. One person who made the call and can defend it.

That clarity, about who decided and why, is what governance actually looks like. Everything else is just process.

In Practice: AI in the Enterprise | Day 1: The Moment Your AI Strategy Became a Governance Problem

I watched it happen in three different organizations within a year. Each one had done the hard work—hired the right talent, built capable systems, deployed models that worked. Then came the moment when someone in the board room asked a question that sounded simple: “Who approved this?”

The room went quiet. Not because they didn’t have an answer. They had too many answers, all contradicting each other.

In one case, the answer was the VP of Analytics. In another, Product. In a third, it was “well, the data science team built it, and operations deployed it, so…” That trailing off matters. It’s the sound of governance architecture breaking under the weight of something genuinely new.

This isn’t a technology problem. It’s an organizational design problem. And most enterprises haven’t yet realized their governance structures—the things they’ve spent decades perfecting for traditional software, for regulated processes, for operational risk—are fundamentally misaligned with how AI actually works in practice.

Where Traditional Structures Break

Traditional governance works because it assumes clear ownership boundaries. The application owner is responsible. The security team reviews. Compliance signs off. Audit checks the box. There’s a chain of custody.

AI breaks this. Not because AI is magic, but because it distributes responsibility across domains that don’t typically talk to each other. A model’s behavior depends on data quality (traditionally data engineering’s problem), training methodology (data science), business logic interpretation (product), deployment infrastructure (operations), and ongoing performance monitoring (analytics or sometimes risk management).

When something goes wrong—a model drifts, produces unfair outcomes, or makes a costly decision—the question “who owns this?” becomes genuinely difficult to answer. Was it the data scientists who built it? The engineering team that deployed it? The business stakeholders who set the success metrics? The person who defined what “fair” means?

I’ve heard CFOs push back on AI initiatives not because they doubt the technology works, but because they can’t see the accountability chain. That’s not caution. That’s competence. They’re asking the right question, just about the wrong structure.

What Governance Actually Needs to Answer

Real governance has to answer three things:

First: Who can approve the deployment decision? Not who built it—who actually says “yes, this goes to production.” The temptation is to assign this to the most senior technical person in the room. That’s backward. This decision is business risk, not technical risk. A model can be technically sound and still a bad business decision. The person who signs off on deployment needs to understand what it does, what can go wrong, and what the costs are. They need to be positioned in the organization so that cost falls on them if it materializes.

Second: Who monitors for the specific failures that matter? Traditional systems have ops teams watching for downtime. But AI systems that are running perfectly fine technically can still be producing systematically biased outputs, or slowly drifting away from the decision quality they had on day one. You need someone looking for those failures. The person has to understand what they’re looking for. And they have to have the authority to pull the cord if they find it.

Third: Who decides what the model is actually supposed to optimize for? This is the one that trips people up. Data scientists are trained to optimize for mathematical objectives—accuracy, AUC, F1 scores. But a business decision that’s technically accurate can still be wrong. A lending model might predict default probability accurately but produce disparate impact. A hiring model might predict job tenure accurately but systematically screen out qualified candidates from certain groups. The technical metrics don’t capture what actually matters.

Getting this wrong creates a specific kind of organizational failure: governance theater. You’ll put in a review process. You’ll create a checklist. You’ll assign AI governance committees. And then you’ll still have decision-making happening in the gaps—skipped reviews, reinterpreted policies, informal approvals. This happens not because people are cutting corners; it’s because the formal structure doesn’t actually address the real decision point.

The Three Gaps Most Organizations Face

In the organizations I’ve watched go through this, three patterns appear consistently.

The first is accountability without authority. Risk and compliance teams are asked to govern AI but don’t have the budget authority, the technical knowledge base, or the position in the decision flow to actually prevent bad deployments. They review after decisions are made. They’re auditors, not governors.

The second is speed pressure meeting governance design. You’ll see this as “we need to govern responsibly, but we can’t slow down.” So you design governance that theoretically makes sense but requires five review gates, each with different stakeholders, none of whom are empowered to make the final call. Then the organization learns to route around it. You end up with faster, less governed decision-making, not slower, more governed decision-making.

The third is confusing process with governance. You’ll see organizations build elaborate approval workflows for AI—more elaborate than they have for traditional software—thinking that more process equals better governance. It doesn’t. Better governance is clearer accountability, better information to the person making the decision, and real authority to say no.

What Actually Works

The organizations that get this right share something: they’ve explicitly designed governance around how decisions actually get made, not around what they wish would happen.

They assign clear approval authority—often to a business owner, not a technical person—with explicit responsibility for the downstream impact if the decision goes wrong. They build monitoring into the system itself, not as an afterthought, with someone empowered to act on what’s learned. They translate business requirements into the actual constraints that matter for model behavior, and make those explicit at training time, not as a risk to be managed after deployment.

None of this requires new technology. It requires thinking through organizational design with the same rigor you’d apply to any other high-risk process.

The moment your AI strategy became a governance problem wasn’t when you deployed your first model. It was when you assumed your existing governance structure would work for something fundamentally distributed across your organization. Most enterprises haven’t yet recognized this moment. When you see the room go quiet at the question “who approved this?”—that’s when you know you’re there.

The good news: this is solvable. It’s just not a technical problem.

Scaling AI FinOps | Lesson 25: Field Notes from the Canopy

The second dry season ended with everybody in the Canopy, and a number on the table that was larger than the one that started all of this.

That is the ending, and I want to be honest that it is the ending, because the version where the bill goes down is a more satisfying story and it is not the one that happens.

The jungle was spending more than it had two years earlier. Considerably more. It was also serving four times the volume of work at a fraction of the cost per unit, retiring things routinely, funding capabilities that continued rather than projects that ended, and able to say, for every part of that number, what it had bought.

The Crow read the pack. It contained the ledger, the variance between claimed and realized, the decay assumptions with their review dates, the unit economics by capability, and a list of four things that had been switched off since the last review.

She signed it.

The Crocodile, who had spent two years opening one eye to say something correct, read it through and found nothing to correct. So he did not open his eye at all, which in this jungle is the highest compliment available.

The Fox noticed. Nobody else did.

If you are starting here

A fair number of people will read this post first, and it is written to work on its own.

The argument of the whole series is one sentence. AI is the first major category of enterprise spend where the cost number, by itself, tells you nothing about whether the money was well spent, because growth in spend is what success looks like and it is also what a runaway failure looks like, and the two are indistinguishable on a bill.

Which means the question people reach for, are we paying too much, cannot be answered. It is not difficult. It is malformed. The answerable question is whether you are getting enough for what you pay, and that requires a denominator, and almost nobody has one.

Everything below follows from that.

The twenty five

The Jungle Floor

  1. A cloud bill tells you what you bought. An AI bill tells you what you did.
  2. A disagreement you cannot resolve is usually a translation failure. Agree what the number counts before you argue about it.
  3. Inference is the layer you budgeted and rarely the layer that hurts.
  4. Shadow AI is not defiance, it is an unfunded requirement with a credit card.
  5. A hundred experiments is not coverage. It is the same experiment funded a hundred times by people who have not met.

Learning to Count

  1. Count what was accepted, not what was produced.
  2. Decide who funds the shared foundation before you decide how to split the water.
  3. Time saved is not money saved until somebody does something specific with the time.
  4. You cannot measure a change you did not measure before.
  5. Write down what you promised, and never edit the entry.

Architecture Is a Financial Decision

  1. Two cost curves, not two products. The only questions are where they cross and how much you trust the volume.
  2. Most of what you send to the cleverest animal does not need the cleverest animal.
  3. You are billed for the question as well as the answer, and the question is usually bigger.
  4. Price the constraint as a design input rather than arguing about it as a tax.
  5. Autonomy is where you stop buying answers and start buying attempts. Budget the tail, not the average.

The Operating Model

  1. A project is funded to end and a capability is funded to continue. Only one gets cheaper the twentieth time.
  2. Fund persistently, reallocate quarterly, and cut something visible in the first year.
  3. Cost and quality cannot be governed in separate rooms.
  4. A ceiling lands hardest on whoever is using the thing most.
  5. A commitment locks your volume, your price, and quietly your capability tier.

The Long Game

  1. Nothing you optimized stays optimized.
  2. Write the sunset criteria at launch, while nobody is attached to it yet.
  3. Assurance is what you pay for the right to change anything quickly.
  4. Spend on literacy rather than on tooling that answers questions nobody knows how to ask.
  5. You cannot manage what you cannot divide.

Four stages

Deliberately behavioral rather than tooling-based, because tooling maturity and actual maturity are only loosely related and the second one is what matters.

Blind. Spend visible in aggregate only. No denominator, no allocation, no acceptance data. Every discussion is anecdote and gets resolved by seniority. Most organizations are here and do not know it, because they have a dashboard.

Attributed. Spend allocated to teams and capabilities. Denominators defined and dated. Showback running. Discussions become factual. Decisions are still slow, because the evidence exists but nothing is built to act on it.

Managed. Unit economics trending. Value ledger live and append-only. Reallocation on a published cadence with real authority. Optimization owned by engineering. The organization can answer whether it is getting enough for what it pays, which is the threshold question.

Compounding. Cost per unit falling while volume grows. Retirement happening routinely rather than during reorganizations. Architecture decisions made on unit economics as a matter of habit. Rare, and realistic within about three years for an organization that starts properly.

Where you actually are

Five questions. Ten minutes in a leadership meeting. Answer honestly rather than aspirationally, which is harder than it sounds.

  1. Can anyone tell you last month’s AI spend broken down by business purpose, not by vendor or model?
  2. For your largest capability, what is the cost per accepted unit of work, and who owns that number?
  3. When was the baseline for your largest benefit claim captured, and was it before or after go-live?
  4. What did you switch off this year?
  5. Are cost and quality reviewed in the same meeting, by the same people?

Three or more no answers puts you at Blind, whatever your tooling suggests. That is not an indictment. It is where nearly everybody starts, and the jungle in Lesson 1 could not have answered any of the five.

Ninety days

If you want to start on Monday, this is the order I would do it in. Each phase has an owner and produces an artifact, because phases without artifacts are intentions.

Days 1 to 30. Find out where you are.

  • Instrument at request level. Team, use case, business purpose on every call. Three weeks of work now, three quarters later. Owner: engineering. Artifact: tagged traffic.
  • Pick one denominator for one capability, and let the people who own it choose it. Owner: the capability owner. Artifact: one dated sentence.
  • Count the pilots and multiply by an honest fully loaded fixed cost. Owner: finance. Artifact: one number.

Days 31 to 60. Build the spine.

  • Publish the allocation rule before you publish any allocated number. Owner: finance. Artifact: one page.
  • Stand up the value ledger, append-only, in a spreadsheet. Owner: finance. Artifact: the ledger.
  • Capture a baseline before anything else ships, and capture the staggered rollout data you are already generating. Owner: the business. Artifact: a dated measurement.

Days 61 to 90. Prove it is real.

  • Run the first review with cost and quality in the same room. Owner: whoever arbitrates. Artifact: decisions.
  • Cut one thing, visibly. Owner: the portfolio owner. Artifact: a retirement.
  • Publish claimed against realized, including the gap. Owner: finance. Artifact: the variance.

Nine items. None of them require a platform purchase. The most expensive is three weeks of engineering time and the most valuable is probably the last one.

The one thing

If all of this reduces to a single argument, it is this.

AI FinOps is not a cost discipline. It is a measurement discipline that produces cost outcomes, and the distinction is not academic. An organization that treats it as cost control will cut things, quickly and confidently, and will cut the wrong ones, because without a denominator every reduction looks like a saving and the panic cap in Lesson 19 looks like a win for five weeks.

An organization that treats it as measurement will spend a slower and less satisfying first year building a denominator, a ledger, a baseline and an honest conversion path. And then it will make cost decisions that are correct, repeatedly, because it can see what it is doing.

The whole thing is the patient work of finding an honest denominator and then defending it against everybody who would prefer a flattering one. That is it. Get it right and the cost questions start answering themselves. Get it wrong and you will cut the wrong things, quickly, with great confidence, and you will not find out for two years.

Jungle Lesson 25

Twenty four lessons reduce to one. You cannot manage what you cannot divide, and the whole discipline is the patient work of finding an honest denominator and defending it against everyone who would prefer a flattering one.

What comes next

That is the series. Twenty five lessons, two dry seasons, one jungle, and a cast who were all right about something and wrong about something else, which is the only kind of cast worth writing.

I said in Lesson 23 that there was a larger subject sitting underneath all of this, and that it did not fit in a Field Kit. It has been visible at the edge of nearly every one of these posts. The Tortoise’s list. The version that changed behavior without warning. The workflow with no owner sitting inside a month end close. The question of who signed off on a system nobody entirely understands.

Cost is the tractable half of enterprise AI. Governance is the half that decides whether any of it survives contact with a serious question, and it is where I am going next, at proper length, starting next week.

Thank you for reading this one. The jungle is quieter than it was.

If you take one thing from twenty five posts, take the denominator. Everything else in this series is an elaboration of it, and an organization that has one is playing a fundamentally different game from an organization that does not, regardless of what either of them has bought.

Scaling AI FinOps | Lesson 24: Growing the Troop

The Mandrill wanted to hire a Head of AI FinOps.

It was, on every visible measure, the right call. The discipline had proved itself over two years. It had a rhythm, a ledger, an allocation model and a set of controls. It clearly needed an owner. The obvious next step for anything that works in an enterprise is to give it a leader, a team, a budget and a reporting line, and the Mandrill had reached that conclusion the way any competent executive would.

The Crocodile talked him out of it in about six sentences.

“I have watched this happen four times,” he said. “Capacity planning. Quality. Security. Cloud cost. Each one worked when it was distributed, so somebody centralized it, and each one turned into a function that produced excellent reports about decisions it was not in the room for. Then it got cut, and everyone said the discipline had not delivered. The discipline delivered. It was moved somewhere it could not reach anything.”

The Mandrill thought about it for a moment and changed his mind, out loud, in front of the room, which is the most senior thing anybody does in this entire story and considerably rarer than the reorganization the Fox proposed eight lessons earlier.

Why centralizing kills it

The mechanism is simple and it is the same every time.

The decisions that move the number are made by engineers choosing routing and context structure, and by product owners choosing scope and sufficiency. Those are the levers. Everything in Act III and half of Act IV lives in the hands of people doing the work.

A central team owns reports. It can describe what happened. It can identify where the money went. It cannot change a routing rule, cannot define an acceptance criterion, cannot decide a quality floor and cannot say no to a use case.

So you get a function with total visibility and no authority, producing increasingly sophisticated analysis of decisions being made elsewhere by people who have not read it. That is not a failure of the people in the function. They are usually good and they are usually frustrated. It is a structural mismatch between where the information sits and where the levers are.

The distributed model

A small central function and embedded accountability everywhere else.

The center owns standards, the ledger, shared tooling and arbitration. It defines how measurement works so that numbers are comparable across the estate, curates the ledger so it stays honest, and resolves disputes that teams cannot resolve themselves.

The capability teams own the decisions. Routing, context, scope, sufficiency, denominators, acceptance definitions, everything that actually changes the number.

Center sets the how of measurement. Teams own the what of decisions. If your center is making decisions, it is too big. If your teams cannot compare their numbers to each other, it is too small.

Four capabilities, not four roles

These have to exist somewhere in your organization. Deliberately described as capabilities rather than headcount, because the fastest way to get this wrong is to read them as a hiring plan.

Unit economics literacy in engineering. The people choosing architecture need to understand the cost consequence at the moment they are choosing, not in a review three weeks later. An engineer who knows what a cache miss costs makes a different decision from one who does not, and no governance process substitutes for that.

Technical literacy in finance. The Crow needs to know what a cascade is. Not to design one. To ask about it. A finance function that can ask what proportion of traffic goes to the top tier is worth more than any dashboard that would have answered the question unasked.

Value definition in the business. Denominators and acceptance criteria. This cannot be delegated to anyone else and it will be invented on your behalf if you leave it vacant.

Arbitration somewhere neutral. Somebody has to resolve cost against quality without owning either side. This is the one role that genuinely does need to sit in the center, and it needs to be senior enough that its decisions stick.

Literacy is the actual lever

The most useful thing in this post, and it is cheap, which is why it gets skipped in favor of something expensive.

Most organizations buy tooling to compensate for a literacy gap. The reasoning is reasonable: our finance team cannot interrogate this, so let us buy a platform that presents it in terms they understand.

What actually happens is that you have bought something that answers questions nobody knows how to ask. The platform is fine. The dashboards are good. Nobody in the room has the vocabulary to look at them and know which follow-up question would be devastating.

A finance team that understands the cost stack from Lesson 3 and the routing question from Lesson 12 asks better questions than any tool answers. That literacy costs a few days of somebody’s time and it compounds permanently. The tool costs considerably more and depreciates.

Invest in the literacy first. If you still need the tool afterwards, you will at least know what you need it for, which is a much better position from which to buy anything.

Roles to be skeptical of

Three job descriptions that get written and should usually not be filled as written.

The AI FinOps Analyst who only reports. Produces the monthly pack, attends the reviews, changes nothing. Frequently a capable person who will leave within eighteen months because they can see the problem more clearly than anyone.

The Cost Optimization Team with no authority over architecture. Recommends. Does not decide. Watches its recommendations get deprioritized against feature work every quarter, forever.

The Center of Excellence that excellent people leave. A perfectly good idea that becomes a holding pattern, because sitting adjacent to the work is less interesting than doing it, and the best people notice that first.

Career design

The structural point, which mirrors Lesson 22 exactly.

Nobody’s career currently rewards this work. There is no promotion attached to a cost per unit that fell. There is no recognition for the routing change that saved a great deal of money and produced no visible output.

So put it in objectives. Cost per unit as a named responsibility in engineering objectives. Denominator ownership in product objectives. Retirements alongside launches in portfolio objectives.

Until it is somebody’s actual job, described in the document that determines their progression, it will remain everybody’s concern and nobody’s work, and it will be done in the gaps by whoever happens to care. That works for about a year and it does not survive that person changing roles.

Three ways this goes wrong

The reporting function. Central team, beautiful outputs, no authority, no change, quietly disbanded in year three with the conclusion that the discipline did not deliver.

Tooling as a substitute for literacy. An expensive platform bought to answer questions the organization has not learned to ask, which it will therefore continue not to ask, at a higher cost.

Accountability without authority. Somebody named responsible for cost who cannot influence architecture, scope or funding. This is the most demoralizing job in enterprise technology and it is created constantly with the best of intentions.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, invest in technical literacy in your own team before you buy anything. Two days of your finance team learning the cost stack will change more meetings than any platform you could procure this year.

If you sit in the Crocodile’s chair, put cost per unit into engineering objectives. It changes behavior faster than any review forum, because it changes what people think about while they are making the decision rather than afterwards.

If you sit in the Mandrill’s chair, resist building the function. Distribute the accountability, keep the center small and standards-focused, and give it arbitration authority rather than delivery responsibility.

For everyone: name who arbitrates cost against quality. If the answer is nobody, then the loudest person in the room is arbitrating, and they are not doing it well because it is not their job and they do not know they are doing it.

Jungle Lesson 24

The decisions that move the number are made by the animals choosing the architecture and the animals choosing the scope, so a central team that owns the reports and none of the decisions will produce beautiful reports and no change. Keep the center small, distribute the accountability, and spend the money on literacy rather than on tooling that answers questions nobody knows how to ask.

Next time: the last one. Second dry season ending, the whole cast in the Canopy, and a bill that is larger than it was at the start of all this. The difference is that this time everybody in the room can explain why. Lesson 25 assembles all twenty five lessons, a maturity model, and a ninety day plan for anybody starting from where the jungle was two years ago.

Changing your mind in public, at seniority, is the rarest and most valuable behavior in this entire story. No framework produces it and every framework depends on it.

Scaling AI FinOps | Lesson 23: The Cost of Trust

The Tortoise and the Fox presented together, which nobody in the jungle would have predicted two years earlier and which was, by that point, the least surprising thing in the room.

What they were asking for was a budget line. Not a control, not a policy, not a gate. A line, sized as a proportion of capability spend, for evaluation infrastructure, ground truth maintenance, output monitoring and incident response capacity.

The Mandrill asked the obvious question, which was why this was not simply part of the compliance budget, given that compliance was clearly what it was for.

The Fox answered it, which was the interesting part.

“Because it is not for you,” he said. “It is for me. Every cost improvement we made last year, all of it, every routing change and every prompt we trimmed, was only possible because we could measure whether we had broken anything. Without that, nobody would have let me near the configuration. The evaluation infrastructure is not what slows us down. It is the only reason we were allowed to move at all.”

The Tortoise, who had been making a version of this argument for three years and getting nowhere, let him make it. She had learned by then that the same sentence lands differently depending on who says it, which is not fair and is entirely true.

Assurance is a cost of production

The framing decides everything downstream, so it is worth getting right at the start.

Frame assurance as a cost of compliance and it becomes a tax. It gets minimized, resented, cut first in a difficult quarter, and owned by a function with no ability to defend it commercially. Everybody agrees it is important in exactly the way that everybody agrees things are important shortly before defunding them.

Frame it as a cost of production and it becomes a line item like any other. This is what it costs to run this capability at a standard where it is allowed to touch anything that matters. Not optional, not virtuous, just part of the cost of doing the thing.

The second framing survives budget season. The first one does not, and I have watched it not survive several times.

What actually costs money

Five components, and the second is the one everyone underestimates.

Evaluation infrastructure. Test sets, harnesses, scoring, and the compute to run all of it. Scales with release frequency, which means it grows exactly as your team gets better at shipping.

Ground truth creation and maintenance. The largest of the five and the one nobody budgets. Human-labeled reference data, produced by people who understand the domain, which decays as the domain moves. This is not a one-time dataset build. It is a standing maintenance obligation and it needs to be funded as one.

Continuous monitoring. Sampling production output and having somebody look at it. A standing labor cost, permanently, not a project.

Incident response capacity. People available to investigate when something goes wrong. This is a retainer rather than a project, and the cost is the availability rather than the usage.

Documentation and audit readiness. Cheap if captured continuously as a by-product of doing the work. Extremely expensive if reconstructed later under time pressure by people who were not there.

Proportionality is the whole game

The single highest-leverage decision in this area, and the one most organizations get wrong in the same direction.

Assurance spend should scale with consequence, not uniformly. A uniform standard applied across the estate overspends dramatically on low-consequence workloads and, almost always, still underspends where it genuinely matters, because the uniform standard was set somewhere in the middle to be politically acceptable.

Tier by consequence. What happens when this is wrong. Who is affected, how quickly is it caught, how reversible is it. A capability that drafts internal summaries reviewed by a person before use is a different object from one whose output reaches a customer unmediated, and treating them identically is not rigor, it is an absence of judgment.

This is the same argument as Lesson 12’s sufficiency and Lesson 9’s evidence grading, applied to a third domain. There is a pattern across this whole series and it is: match the intensity of the thing to the stakes of the thing, and resist the temptation to apply one standard everywhere because it is easier to explain.

The economics of getting it wrong

Three costs when something goes wrong, and the third is the largest and appears in no risk register I have ever seen.

The first is remediation. Fixing it, notifying whoever needs notifying, correcting whatever went out. Visible and usually bounded.

The second is the freeze. While an investigation runs, everything else stops. Every other capability gets a second look. Deployments pause. This is expensive and it is at least predictable.

The third is the durable slowdown afterward. New controls, extra approval layers, a review board that did not exist last month, and a general institutional caution that persists for a year or more after the incident and applies to everything, including all the things that were working perfectly.

That third cost typically dwarfs the first two combined. It is also the one that never gets quantified, because it arrives gradually, is distributed across every team, and looks like prudence rather than cost.

Assurance is what buys speed

The argument the Fox made, and the reason this lesson belongs in a series about money rather than in one about risk.

Every cost optimization in Act III is a potential quality change. Routing to a cheaper tier. Trimming context. Reducing agent steps. Every single one of those is a change to behavior, made in pursuit of a lower number.

An organization that cannot measure whether it has broken anything will not permit those changes, and it is right not to. So Act III becomes theoretical, the routing stays as it was configured during the pilot, and the largest available cost lever in your estate remains untouched, permanently, for reasons that will never be written down anywhere.

Assurance infrastructure is what converts optimization from a gamble into a decision. That is not a compliance benefit. That is the enabling condition for the entire third act of this series, and it should be argued for on exactly those terms in front of exactly the people who control the budget.

A larger subject than this series

I want to be clear about the limits of what I have covered here, because this lesson has been the closest I have come to a domain that deserves considerably more than one post.

Everything above treats assurance as an input to a cost decision. That is a legitimate framing and it is the one this series needed. It is also a narrow slice of a much larger question about how enterprises govern systems whose behavior they cannot fully specify, who is accountable when those systems are wrong, and what it means to approve something you do not entirely understand.

That question is the one I keep running into at the edge of every engagement, and it is the one I intend to write about next, at proper length, because it does not fit in a Field Kit.

Three ways this goes wrong

Assurance as a gate rather than a capability. A review board with no infrastructure underneath it, producing delay without evidence. This is the most common shape and it is the most expensive, because it costs time and buys nothing.

Uniform standards. The same rigor everywhere, overspending on the trivial, underspending on the consequential, defended on the grounds of consistency.

Retrofitted documentation. Nothing captured as you go, then a frantic and enormously expensive reconstruction by people working from memory when somebody finally asks.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, put assurance on its own budget line, sized as a proportion of capability spend and tiered by consequence. Buried inside a project budget it is the first thing cut and the last thing anyone defends.

If you sit in the Crocodile’s chair, automate evidence capture into the pipeline. Continuous is cheap and retrospective is not, and the gap between those two costs is larger than almost anybody expects until they have lived through it once.

If you sit in the Mandrill’s chair, classify your use cases by consequence honestly. Only you can do this. Over-classifying is expensive and under-classifying is worse, and nobody else in the organization has the standing to make the call.

For everyone: budget ground truth maintenance as recurring rather than one-off. It decays, it underpins everything else, and it is the line that gets cut in year two by somebody who thinks the dataset is finished.

Jungle Lesson 23

Assurance is not what you pay to satisfy the tortoise, it is what you pay for the right to change anything quickly. Size it by consequence rather than uniformly, budget the ground truth as a recurring cost because it decays, and remember that the most expensive part of an incident is not the incident, it is the eighteen months of caution that follow it.

Next time: the Mandrill decides to hire a Head of AI FinOps and build a proper function, which is the obvious move and the wrong one. The Crocodile talks him out of it in about six sentences, and the Mandrill changes his mind in public, which is the most senior thing anybody does in this entire story. Lesson 24 is about roles, skills, and why this should not be a team.

The reframe in this piece is worth stealing whoever you are. Assurance argued as compliance gets minimized. Assurance argued as the condition for being allowed to move quickly gets funded. Same infrastructure, same cost, completely different meeting.

Scaling AI FinOps | Lesson 22: Nobody Gets Promoted for Switching Things Off

The last Hummingbird was switched off on a Thursday afternoon, twenty one lessons after it first appeared, and it took about four minutes.

Nobody could establish what it had been for. It had a name that suggested a purpose, but the purpose it suggested had been served by something else for over a year. The animal who built it had moved to another territory two seasons earlier and, when asked, remembered starting it and did not remember why.

It had been running the entire time. It had consumed a small amount of money every single day for two years, quietly, correctly, producing output that went into a location nobody had opened.

The Fox switched it off and waited a fortnight, which was the agreed protocol by that point.

Nothing happened. No complaint, no escalation, no broken workflow. After the fortnight it was removed properly.

The Hyena, who had made a good deal of noise about the Hummingbirds over the preceding two years, was oddly quiet about this one. “It ran for two years,” she said. “Somebody built it because they cared about something. And in the end nobody noticed it stop.”

Which is roughly the correct emotional register for this topic, and the reason it is the least discussed subject in the whole discipline.

The incentive problem, stated plainly

Launching is visible, credited and promotable. There is a demonstration, an announcement, a slide with your name on it and a line in your performance review.

Retiring is invisible, thankless, and carries all of the risk. Nobody is congratulated for it. The upside is a cost reduction that will be attributed to the portfolio rather than to you. The downside is that you broke something, and that will be attributed to you very specifically.

So every enterprise accumulates. This is not a discipline failure and it is not a maturity failure. It is a rational response to an incentive structure, and it will not be fixed by encouragement, by policy, or by anyone writing the word rationalization on a slide. It gets fixed by changing what is rewarded, and by nothing else.

What accumulates in an AI estate

Five categories, in roughly ascending order of how much they cost and how invisible they are.

  • Orphaned pilots. The Lesson 5 population. Individually small, collectively significant, and completely undefended once anybody counts them.
  • Superseded capabilities. Something better was built. The old thing still serves eleven people who never migrated. It costs a full capability’s fixed overhead to serve eleven people.
  • Evaluation harnesses for retired models. Still running, still scheduled, still consuming, testing something that no longer exists.
  • Duplicate capabilities. Two territories built the same thing before the portfolio function existed. Both work. Neither owner will volunteer to be the one that stops.
  • Retrieval indexes nobody queries. The purest case, and the most expensive of the five.

The index nobody reads

The last one deserves its own section because it is the cleanest illustration of why AI waste is different from every other kind, which was the argument that opened this entire series.

A retrieval index costs money every day to keep current. It re-processes content as content changes. It does this whether or not anybody has ever queried it. It appears on no dashboard as waste, because it is running perfectly, doing exactly what it was built to do, without error.

And there is no signal anywhere in your systems that says nobody has asked this anything in fourteen months. Unless somebody specifically goes looking for query counts against index cost, it is invisible forever, and it costs more than any of the orphaned pilots.

This is the Lesson 1 point arriving at its final form. Waste that looks exactly like work, running correctly, indefinitely, because the only thing that would reveal it is a question nobody has thought to ask.

Sunset criteria, written at launch

The intervention that actually works, and the reason it works is timing.

Three things, written into the launch approval, before anybody is attached to anything.

A usage floor. Below this many units per period, this gets reviewed for retirement. Not automatically killed, reviewed, which is a much easier thing to agree to in advance.

A value floor. Below this contribution against the ledger, same.

A named alternative. If this is retired, what do its users do instead. Answering this at launch is trivial. Answering it three years later requires an investigation.

The timing is the entire trick. At launch, nobody is emotionally invested and the criteria feel like sensible hygiene. Three years later, the same conversation is a negotiation with somebody defending their creation, and nobody can objectively assess their own work at end of life. The criteria have to predate the attachment.

Dark launch the shutdown

The single most useful and most under-used technique in this post.

Before removing anything, turn it off quietly and wait. Two weeks is usually enough. Do not announce it. Do not send a consultation email, because a consultation email guarantees that somebody who has not used it in a year will explain that they might need it.

Then see who complains.

The proportion of the time that nobody complains is genuinely surprising the first few times you do it. And when somebody does complain, you have learned something specific and valuable, which is that this thing has exactly one real dependency and you now know who it is and can talk to them directly.

The formal version of the process is identify, notify, migrate, dark launch, remove. But the dark launch step is what turns decommissioning from a negotiation into an observation, and observations are much easier to act on than opinions.

Making retirement promotable

The structural fix, and it is the only one that lasts.

Report decommissioning as a headline achievement, in the same document, with the same prominence as launches. Count capabilities retired alongside capabilities launched, as a matched pair, so a portfolio with twelve launches and zero retirements looks like what it is.

Give the portfolio owner from Lesson 16 explicit credit for reductions. Put it in objectives. Make it the thing somebody is measured on rather than the thing everybody agrees is important.

Organizations that do this find that things get switched off. Organizations that merely encourage it find that nothing does, and then conclude that their people lack discipline, when what their people lack is a reason.

The archaeology problem

One last point, which is really a Lesson 6 problem arriving late.

After eighteen months nobody remembers why something exists. The person who knew has moved. The document that explained it was in a folder that got reorganized. So the decision to keep it defaults to yes, because keeping something you do not understand feels safer than removing something you do not understand.

Documented purpose at launch is what makes retirement possible later. One sentence, written when it was obvious, saying what unit of work this exists to improve and how you would know. That sentence is what a future colleague uses to decide, and its absence is why the Hummingbird ran for two years.

Three ways this goes wrong

Retirement by attrition. Waiting for things to become obviously useless, which they never quite do, because something running correctly never looks useless.

The hostage user. One team still depends on it, so it stays forever at full cost serving four people, and nobody ever does the arithmetic on what those four people are costing.

No sunset criteria. Every retirement becomes a fresh negotiation with somebody emotionally invested, which means retirements happen only during reorganizations, which is not a strategy.

The Field Kit

Concrete things to do this week.

If you sit in the Crow’s chair, ask for the list of capabilities retired this year. If the answer is none, your portfolio is growing monotonically and your cost curve is doing the same thing whether or not anyone has noticed.

If you sit in the Crocodile’s chair, find what is still running with no traffic and no owner. Start with the retrieval indexes and compare index cost against query counts. That single comparison finds more money than most optimization projects.

If you sit in the Mandrill’s chair, write sunset criteria into your launch approvals. This is the only moment anybody will agree to them and it takes ten minutes at a point when nobody objects.

For everyone: dark launch the shutdown. Turn it off, wait a fortnight, see who notices. It is remarkable how often the answer is nobody, and it is the cheapest experiment available to you.

Jungle Lesson 22

Every enterprise is structurally incapable of switching things off, because launching is promotable and retiring is invisible. Write the sunset criteria at launch while nobody is attached to it yet, count retirements the way you count launches, and when in doubt, turn it off quietly and see who complains. Surprisingly often, nobody does.

Next time: the Tortoise and the Fox present together, which nobody in the jungle would have predicted two years earlier, and make a case that assurance is not what you pay to satisfy compliance but what you pay for the right to change anything quickly. Lesson 23 is about the cost of trust.

Every organization has its own Hummingbird, running correctly, costing a little every day, serving a purpose nobody remembers. Finding it is not hard. Deciding to look is the hard part, because nobody has ever been thanked for it.