In Practice: AI in the Enterprise | Day 16: Governance by Committee vs. Governance by Design: Why Your AIOC Might Be Theater

I walked into a Fortune 500 company’s AI governance meeting and saw the org chart: 14 people, representatives from Legal, Compliance, Risk, Product, Engineering, Finance, and three layers of leadership. The meeting itself had run for six months with no shipping decisions made, but plenty of documentation produced. When I asked the head of the program what problem this committee was actually solving, she said: “We’re ensuring governance.”

That’s not governance. That’s the appearance of governance.

Governance is a decision-making framework. The best decision-making frameworks are not democratic. They’re not designed to include everyone. They’re designed to ensure the right people are involved at the right decision points, with clear authority, clear escalation paths, and clear consequences for being wrong.

The difference between these two things is the difference between a committee that produces theater and a governance structure that produces decisions.

How Committees Become Theater

Here’s what happens when you design governance through consensus:

You assemble people from different functions with different incentives. Finance wants to minimize risk and cost. Legal wants comprehensive documentation. Product wants speed. Engineering wants stability. They all get a vote.

The only outcome that keeps everyone happy is the one that offends everyone equally. So you set rules that are both conservative (Legal is satisfied) and bureaucratic (Finance is satisfied) and slow (Risk is satisfied) and documented (Compliance is satisfied). Everyone signs off. Nobody is happy. The system is designed not to make decisions, but to distribute responsibility so that when something goes wrong, you can point to the process instead.

The meetings get longer. More people get invited. Someone suggests a charter document. Someone else suggests subcommittees. Pretty soon you have a governance structure that exists to ensure it can explain why it didn’t make the decision anyone wanted.

I’ve watched this happen in companies with smart people making rational decisions. They’re not trying to build theater. They’re just following the logic of “if we want everyone to agree, we need to make decisions nobody disagrees with,” which means decisions that satisfy nobody.

The problem is that AI governance is not primarily a legal or compliance problem. It’s a speed and risk tradeoff. Committees are designed to minimize disagreement, not to optimize for speed or actually manage risk.

What Governance by Design Actually Looks Like

The companies that are handling this well have made a structural decision: they’ve separated the decision into parts.

Who decides what gets built: Product and Engineering. Speed is necessary here. The committee has no role. They have constraints (legal, compliance, risk), but constraints are not votes.

Who decides if those constraints are met: Legal, Compliance, Risk. Their job is not to vote. Their job is to flag when a proposal violates the constraints. If the constraint is “all customer data must be encrypted in transit,” then a proposal that doesn’t encrypt in transit gets sent back. The decision-maker has to either change the proposal or change the constraint (which escalates).

Who has authority to change constraints: Finance and executive leadership. This is the only place where you need consensus, because changing the constraints is actually changing the business model or the risk appetite. If the constraint is “all AI systems must use approved models,” and Product wants to use an unapproved model, that’s a constraint change. That’s the conversation for leadership.

Who owns the outcome: Product and Engineering, with escalation to leadership if constraints are violated. This is critical. The people who make the decision have to own the consequences. If they ship something that violates compliance, it’s on them, not on “the committee.”

This structure has a critical property: it makes decisions fast because there’s no voting. It’s also clear about what’s being constrained and why. And it’s clear who’s accountable.

Why This Actually Reduces Risk Better Than Committees Do

Here’s the counterintuitive part: decentralized governance with clear ownership actually manages risk better than consensus governance.

When everyone is responsible, nobody is. The committee that approved something problematic can always point to the process. But when Product and Engineering own the decision, and they know that Legal will flag constraint violations, and they know that constraint violations escalate, they think more carefully about what they’re shipping.

Also, a committee of 14 people will never move as fast as a product team. So you actually get worse outcomes: slower decisions and less careful consideration because the decision-making is so slow that nobody’s paying attention anymore.

The framework I’m describing isn’t “let engineering do whatever they want.” It’s “engineering makes the call, but within clearly defined constraints, with clear escalation paths when those constraints conflict.” That’s radically different from consensus.

The Conversation About What You’re Actually Trying to Prevent

Here’s where most companies go wrong. They build a governance committee without first deciding what they’re actually trying to prevent.

Are you trying to prevent: – Systems that break laws (you need Legal with veto authority, but that’s narrow) – Systems that harm customers (you need Product with clear accountability) – Systems that lose money (you need Finance with clear cost constraints) – Systems that offend people (this is a different problem; it’s about policy, not governance) – Systems that break architecturally (you need Engineering with clear technical standards)

The answer matters because different things need different governance structures. If you’re trying to prevent legal violations, you need a clear compliance rule and someone with authority to enforce it. If you’re trying to prevent customer harm, you need clear product standards and someone accountable for them. If you’re trying to prevent everything, you’ve got a committee.

Most enterprises build committees because they’re trying to prevent everything, and they haven’t had the conversation about what the actual constraints are. So they generate rules to cover every possible scenario, which means nothing ships, which means the governance structure prevents you from building AI at all.

What Actually Changes When You Design for Decision-Making

The shift from committee to designed governance requires:

  1. Clarity about constraints, not votes. What are the actual constraints? Make them explicit. “All AI systems must use approved models” is a constraint. “The committee must agree” is not. Constraints are narrower and measurable.

  2. Clear escalation paths. When a constraint conflicts with a business goal, where does that decision go? Not “back to the committee.” To a specific person or leadership group with actual authority to change the constraint.

  3. Accountability that sticks. If Engineering ships something that violates a constraint, it’s on them (with escalation to leadership if it’s a material risk). This makes them think carefully about what they’re shipping.

  4. Separation of concerns. The people who decide if you build something (Product) are different from the people who decide if you can build it that way (Legal/Compliance). Both groups are necessary. Neither has a vote.

This requires culture shift. Companies have been trained to believe that shared decision-making is fair, and that consensus is legitimate. Neither is necessarily true. Clear authority with clear accountability is fairer and faster than democratic committees that don’t actually decide anything.

The Governance Structure No One Talks About

Here’s what I never see in organizational documents: “The AI Governance Committee will not make decisions. Its purpose is to flag constraints and escalate conflicts. The decision authority rests with [specific role].”

That sentence would solve most governance theater problems. But it requires a company to say clearly: “Product owns the decision. Legal flags constraint violations. Engineering owns the technical quality. Finance owns the cost constraints. Leadership breaks ties.”

The theater happens when you try to maintain the fiction that everyone is equally responsible for the decision. The real governance happens when you distribute responsibility clearly and make someone actually accountable.

Your AIOC might be fine. But if you can’t answer the question “who actually decides if we build this?” in less than a sentence, it’s probably theater.

In Practice: AI in the Enterprise | Day 15: Strategic Independence in Your AI Stack: What It Actually Requires

A CTO told me they’re building their own fine-tuned version of an open-source model to reduce vendor dependency.

I asked: Why?

The answer was: “So we’re not locked into one vendor.”

That’s the right concern. The execution is usually backwards.

Building your own fine-tuned model doesn’t make you independent. It makes you dependent on a different thing: your ability to maintain it. And that’s a dependency most enterprises underestimate.

Independence Isn’t About Which Model You Use

Strategic independence in AI is a real problem. I’ve written about it already (vendor dependency, lack of abstraction layers). But what independence actually requires is often misunderstood.

Most organizations think independence means: “Don’t use foundation models. Build or fine-tune your own.”

That’s treating the symptom, not the disease.

The disease is structural dependence on any single infrastructure or capability that you can’t replicate if the vendor relationship changes. The cure is designing your systems so that core capabilities either: a) exist across your organization as institutional knowledge, b) can be replaced without rebuilding the entire application, or c) are decoupled from your core business logic.

Building your own fine-tuned model can accomplish this. But so can building abstraction layers. So can designing for model rotation. So can maintaining real technical depth in how your models work.

The question isn’t “should we build or buy.” The question is “what do we need to maintain control over, and how do we actually maintain it?”

What Independence Actually Costs

Let me be concrete about what it means to build and maintain your own fine-tuned models.

First: initial build cost. You need engineers who understand model training, which means data scientists with 5+ years of experience. If you need to build from scratch (not fine-tune), you’re talking $1-3M in development cost, 6-12 months of timeline, and a team of 4-8 people.

Second: ongoing compute costs. A foundation model API costs you based on usage. Running your own model costs you compute (GPU infrastructure, maybe $100K-500K annually depending on scale) plus the electricity, networking, and datacenter costs that come with that.

Third: maintenance burden. When a new model comes out that’s better, should you retrain? When your domain shifts and your training data becomes stale, should you retrain? When you discover bias in your model, how do you fix it? Each of these is an active maintenance project, requiring someone to be responsible for it.

Fourth: talent concentration. You now have a capability that only a few people in your organization understand deeply. If your model expert leaves, or is overloaded, or moves to a different team, you have a problem. You’ve traded vendor dependency for person dependency.

Fifth: opportunity cost. While you’re building and maintaining your own model infrastructure, you’re not building the applications that use it. You’re spending engineering time on infrastructure that, for most organizations, a vendor could provide more efficiently.

The argument for building your own typically goes: “This gives us independence.”

The reality usually is: “This trades one dependency for another, often a more expensive one, and concentrates risk in a different place.”

When Building Your Own Actually Makes Sense

There are scenarios where building your own model is the right answer:

Scenario 1: Proprietary advantage through data. You have data that nobody else has access to, and that data gives you competitive advantage. You’re not just using a foundation model as a starting point. You’re incorporating proprietary data to create a capability that competitors can’t easily replicate. The privacy and competitive protection from keeping this internal is worth the build cost.

Example: A financial services firm with 50 years of internal transaction history that becomes part of the model. A manufacturing firm with sensor data from their equipment that’s unique to their process.

This is a real reason. But notice: you still might use a foundation model as a base and fine-tune it with your proprietary data. You don’t necessarily need to train from scratch.

Scenario 2: Extreme scale or cost sensitivity. You’re running this model so many times that the marginal cost of API calls is expensive relative to running your own infrastructure. You’ve done the math, and your compute costs are 30% lower running your own.

Example: A company processing millions of requests per day where each API call costs money, and running your own infrastructure has better economics.

This is a real reason. But notice: you still need the expertise to maintain the infrastructure. And you need to be comfortable with the operational burden.

Scenario 3: Regulatory or data residency requirements. Your industry requires that models be trained on data that never leaves your infrastructure. You can’t send data to a third-party API. You need the model to run in your environment.

Example: A healthcare provider with patient data, a financial institution with transaction data, a government agency with classified data.

This is a real reason, and it’s absolute. You have no choice. But notice: you might still use a foundation model as a base (open-source, self-hosted, or fine-tuned on your infrastructure) and then add your domain-specific fine-tuning.

In all three scenarios, the justification is about data (proprietary, sensitive, scale-related). Not about “independence from vendors.”

What Independence Actually Requires (If You’re Not Building Your Own)

If you’re not building your own model, but you still want strategic independence, here’s what actually works:

First: Design abstraction layers that are model-agnostic.

Instead of writing code against a specific model vendor’s API, write code against an abstraction. Different models can satisfy that abstraction. It’s more complex. It’s less efficient (you can’t optimize for vendor-specific quirks). But if your primary model changes, you can migrate without rebuilding.

This is the most practical approach for most organizations. It gives you real independence without the cost and complexity of building your own. This applies whether you’re building on any cloud AI platform, open-source models, or proprietary APIs—the principle is architectural, not about any specific vendor.

Second: Maintain real technical depth about how your models work.

Don’t treat models as black boxes. Understand what they’re doing, why they work, what they fail at. This requires people who can debug models, understand prompt behavior, interpret outputs. If you have people who really understand models deeply, you’re not dependent on any single vendor because you have the skill to work with alternatives.

Third: Build model rotation into your roadmap.

Plan for changing models not as a crisis scenario, but as a regular architectural activity. Every 18-24 months, do an evaluation of foundation models. Test alternatives. Understand switching costs. Keep up relationships with multiple vendors. When you do need to switch, you’re not starting from zero.

Fourth: Keep the model separate from your application logic.

The model is a component, not the core. Your application has business logic that’s independent of which foundation model it’s using today or tomorrow. The model is pluggable. This architecture costs more upfront but gives you real flexibility later.

Fifth: Maintain financial optionality.

Budget for alternatives. Set aside 5-10% of your AI budget for evaluating options. Build organizational practices where model choice is a strategic decision, not a habit. If your only model is because it’s the one you started with, that’s dependency. If your primary model is chosen and re-evaluated regularly, that’s strategy.

Why This Matters For Your Board

Strategic independence in AI infrastructure will matter more as enterprise AI becomes more central to business operations. Right now, most organizations are still experimenting. Dependency feels theoretical.

Over the next two years, as AI systems become embedded in core business processes (customer service, decision-making, optimization), a vendor change becomes a significant operational event. Your systems, your teams, your customer-facing capabilities depend on that vendor’s continued cooperation and reasonable terms.

The organizations that will have optionality are the ones that designed for it from the start. Not the ones that panic later and decide to “build independence” by rewriting their infrastructure from scratch.

Building that independence doesn’t require building your own models. It requires:

  • Architecture decisions that allow model swapping
  • Technical depth so you understand models, not just use them
  • Regular evaluation of alternatives so you don’t wake up dependent
  • Clear separation between model and application
  • A serious answer to “what would we do if this vendor changed terms?”

The organizations that get this right won’t necessarily switch models. They’ll just have the ability to, which means they’ll negotiate better, maintain strategic clarity, and adapt faster when the foundation underneath them shifts.

Independence is possible. It just doesn’t come from building your own model. It comes from designing your systems as if you might need to change it.

In Practice: AI in the Enterprise | Day 14: The AI Adoption Curve Nobody Talks About: Why Your Best Tools Get Abandoned

You built something good. The model works. The infrastructure is solid. You’ve got buy-in from the business sponsor. You’ve trained the team. You’re ready to go.

Six months later, 40% of the team is using it. Eighteen months later, it’s been deprecated.

This isn’t a technical problem. It’s not a governance problem. It’s an adoption problem that looks like a technical problem, and most organizations don’t see the pattern until it’s too late.

What Most Organizations Get Wrong About Adoption

The standard narrative about adoption goes like this: “We built a tool. People didn’t want to use it. Maybe the tool wasn’t good enough, or the training wasn’t good enough, or the business case wasn’t compelling enough.”

And then the organization either: a) doubles down on training, b) adds more features to make it more compelling, or c) gives up and builds something else.

None of these is wrong exactly. But they’re all treating adoption as primarily a tool or training problem. It’s not.

Adoption in organizations isn’t like adoption in consumer markets. In consumer markets, if you build something better than alternatives, people use it. They compare options and pick. They vote with their usage.

In enterprises, adoption is a work context problem. People don’t choose what to use based on whether it’s best. They choose based on: Does this fit into how I already work? Does using this create friction with other parts of my job? Does my manager expect me to use this? Do the people I work with use it?

The best tool in the world will fail if it doesn’t fit into the actual context of how the work happens.

The Real Adoption Curve (Not The One In The Literature)

Most adoption literature talks about an S-curve: early adopters, early majority, late majority, laggards. You get 10% adoption, then 20%, then it accelerates. Classic diffusion of innovation.

That’s not what happens with enterprise tools.

What happens with enterprise tools is:

Month 0-2 (Honeymoon): High enthusiasm. The team that built it is excited. The sponsor is excited. Early adopters are excited. You see 50-60% usage among people who could theoretically use it. People are trying it. Everyone’s optimistic.

Month 3-4 (Friction Discovery): Usage starts to flatten. You’ve got 50% usage, and it’s not growing. The people who are using it are hitting edge cases. The tool doesn’t work perfectly for their specific workflow. You ask why usage isn’t higher, and you get feedback like: “It doesn’t integrate with my existing process.” “It’s faster for me to do it the old way and copy the result in.” “I can’t use this when I’m at a client site.” “The output isn’t in the format I need for my downstream system.”

None of these are “the tool doesn’t work.” They’re “the tool doesn’t work in my context.”

Month 5-6 (The Push-Back): Leadership decides usage is too low. They mandate using the tool. Or they add requirements: “You have to use this by the end of Q2.” Or they tie it to performance reviews: “Usage of this tool is part of your evaluation.”

This is when you learn whether you’ve actually solved the context problem or just built a nice tool.

Month 7-12 (Deterioration): If you fixed the context problems, usage stabilizes at a higher level. If you didn’t, usage drops below the mandate level as soon as the mandate is lifted. People find workarounds. They convince themselves the tool is “not quite right for my situation” so they go back to the old way. They batch the work they have to do with the tool with work they do the old way, then copy the results in. Adoption grinds down.

Month 13+ (Abandonment): The tool is still there. Some people use it. But it’s no longer the default. New people aren’t trained on it because “it didn’t really take off.” The organization moves on to the next initiative.

The standard response to this pattern is: “The tool wasn’t good enough.” The actual response should be: “We didn’t understand the context well enough.”

What Context Actually Means

When I say “context,” I don’t mean “people don’t like change.” I mean the specific, structural constraints on how work actually happens.

Say you built an AI model that improves sales forecasting. It’s a good model. Salespeople who use it get better at predicting. But your organization has a specific context:

  • Salespeople are evaluated monthly on their forecast accuracy
  • They’re compensated based on hitting their forecast
  • The model updates twice a week
  • But they submit their forecast once a week
  • And once submitted, changing the forecast requires approval from their manager
  • Who is traveling this week

In this context, the salesperson faces an explicit cost to using your better model: they have to get approval to change a forecast they’ve already submitted. The benefit (better accuracy) is diffuse and doesn’t show up in their evaluation immediately. So they don’t use it.

This isn’t a training problem. This isn’t a “people resist change” problem. This is a structural misalignment between when the tool is useful and when the person is incentivized to use it.

The way you fix this isn’t by training harder. It’s by understanding the structural context and redesigning around it. Maybe the model needs to run on the person’s schedule, not its own. Maybe the forecast needs to be re-submittable without approval. Maybe the evaluation system needs to change. Maybe you need to give people a 30-day look-ahead so they can submit more aggressive forecasts and let the model refine them.

None of this is about the model. It’s all about the work context.

Three Patterns That Kill Adoption

I’ve watched three adoption-killing context problems show up repeatedly:

1. Workflow Mismatch

You built a system that works great if people follow a certain workflow. But that workflow doesn’t match how people actually work.

Example: You built a data preparation tool that works great if people upload data, review it, and run preprocessing. But people’s actual workflow is: I’ve got data in a database, I need to run analysis in a spreadsheet, and I need results in a presentation. Your tool is in the middle of their workflow but doesn’t integrate to the beginning or the end. They have to export from the database to the tool, import the results into the spreadsheet, then copy the numbers into the presentation. That’s three tools instead of two, so they don’t use yours.

2. Incentive Misalignment

You built a system that benefits the organization, but the person using it doesn’t benefit individually.

Example: You built an AI system that optimizes staffing across departments. It’s great for the organization. But a department manager’s bonus is based on their local efficiency, not global efficiency. Using your system might move one of their people to another department, which makes them look bad locally. So they don’t use it even though it’s better for the company.

3. Authority or Credibility Gap

You built a system that provides recommendations or guidance, but the person using it doesn’t trust the system or its recommendations.

Example: You built a model that recommends treatment options for doctors. The model is accurate. But doctors are trained to make their own clinical judgments. They don’t trust a model’s output more than their own experience. Unless you can demonstrate that the model is more accurate than they are—which requires time and data they don’t have—they won’t trust it enough to use it over their existing judgment.

Each of these patterns shows up differently in different contexts, but they all have the same underlying structure: the tool is good, but using it costs the person something (time, authority, credibility, incentive alignment) that makes it rational not to use it.

How To Actually Fix Adoption

The organizations that get adoption right don’t start by building the model. They start by understanding the context.

Map the workflow. How do people actually do this work today? Don’t ask them—watch them. Or better, do the work with them for a few days. You’ll see things they don’t even notice they’re doing.

Identify friction points. Where in the workflow would your tool create extra steps? Where would it require people to change how they work? Where would it require approval or credibility they don’t have yet?

Understand incentives. Is the person who will use this tool evaluated on the thing your tool helps with? Or are they evaluated on something else? If they’re evaluated on something else, using your tool is a cost to them, not a benefit.

Design for the context, not against it. Build the tool to fit into the workflow people actually use, not the workflow you wish they used. If people need results in spreadsheets, make sure your tool outputs to spreadsheets natively. If people are evaluated on local metrics, make sure your tool makes them look good locally, even if it’s making global trade-offs.

Plan for the adoption curve. Don’t expect the S-curve. Expect the friction curve: high, then flat, then mandated, then deteriorating. Plan for each phase. Build in the changes you need to make to get past the friction phase before you deploy.

Why This Matters

The cost of abandoned tools is higher than you think. You spent development time. You spent infrastructure costs. You spent training time. You spent organizational attention. And then six months later, it’s not used.

But the bigger cost is that the organization loses confidence in your AI program. One abandoned AI tool creates skepticism about the next one. The team that built it loses credibility. The sponsor becomes hesitant. By the time you want to try again, you’ve burned through social capital.

The organizations that are winning with AI aren’t the ones building the most sophisticated models. They’re the ones that build models that fit into how people actually work, and then systematically manage the adoption curve until the tool becomes part of the standard workflow.

Adoption isn’t something that happens if you build it right. It’s something you design for, plan for, and manage through the friction phases.

In Practice: AI in the Enterprise | Day 13: Why Your AI Project Budget Blew Up (And How to Prevent It in the Next One)

A VP of Engineering told me about their AI project last year. Budget was $2M. Actual spend was $6.8M. When I asked what happened, she said: “We didn’t understand what we didn’t understand.”

That’s not a failure of planning. That’s a failure of the budgeting framework itself.

The organizations I’ve seen blow through AI budgets aren’t usually incompetent. They’re using budgeting frameworks that were built for a completely different kind of work. And they don’t realize it until six months in, when the budget is half-spent and the project is 20% complete.

Why Traditional Project Budgeting Fails For AI

Traditional software project budgeting assumes a few things: – You know what you’re building – You know roughly how long it will take to build – You know what “done” looks like – You know what success means

AI projects violate every one of these assumptions.

You think you’re building a recommendation system. Once you start, you realize your data quality isn’t good enough. You spend six weeks fixing data. Then you realize your features don’t predict what matters. You spend another eight weeks rebuilding features. Then the model works, but it produces biased outputs. You spend another month on fairness improvements. None of these were “surprises” in the sense that they were unknown risks. They were surprises in the sense that you didn’t know they were going to be the blocking issues.

The fundamental problem is that AI projects are exploration work disguised as execution work.

In traditional software, you explore during the design phase. You do architecture reviews. You write specs. You plan. Then you build. The build phase is well-understood. You have precedent. You have patterns. You have experienced people who know how long it takes.

In AI projects, the exploration doesn’t stop. You’re executing, and you’re discovering simultaneously. You’re building the model and learning what the model can do. You’re testing it and discovering what it fails at. You’re deploying it and discovering what real-world performance looks like. Each phase reveals information that changes the phase that comes next.

So you budget for the plan you made in month one. But month three reveals that the plan was 40% wrong. Do you stop and replan? Most organizations don’t. They replan and continue. Which means you’re now $800K into a $2M budget, and you’ve got $1.2M left to finish work that actually costs $1.8M based on what you know now.

Where The Costs Actually Hide

There are five categories of AI project costs that traditional budgeting frameworks don’t account for:

1. The Data Prep and Exploration Tax

You need training data. In traditional software, data is often already there. In AI, you often need to acquire, clean, label, and prepare data. A team of people might spend 40% of the project doing this. And you don’t really know how much you need until you build the first model and discover what’s missing.

One organization budgeted $200K for data prep. Actual spend: $800K. Why? They discovered their labels were inconsistent. They had to relabel a dataset of 100K examples. They discovered they needed geographic variation they didn’t have. They had to acquire new data. They discovered the time-window for some features was wrong. They had to rebuild the dataset. This isn’t incompetence. This is the normal cost of working with data when you don’t fully understand it upfront.

2. The Model Exploration and Iteration Tax

You build a model. It gets 87% accuracy. Is that good? You don’t know without a baseline. You don’t know without understanding what 87% accuracy means for your business. Does a model that’s 87% accurate but biased against a particular group meet your requirements? You don’t know until you measure it.

So you iterate. You try different architectures. You try different feature sets. You try different training approaches. Each iteration takes computational time and engineer time. A team might do 30-40 iterations before landing on something that’s both technically sound and meets business requirements.

Budget for this? Most organizations don’t. They budget for “build the model.” They don’t budget for “understand the model well enough to deploy it.”

3. The Validation and Governance Overhead

Before you deploy, you need to validate that the model does what you think it does. You need to test it against edge cases. You need to check it for bias. You need to make sure it handles distribution shift. You need to document the failure modes.

This is where governance overhead becomes a cost item, not a process item. A data scientist might spend 30% of their time building the model and 70% of their time validating, documenting, and preparing for governance review.

Most organizations budget for “build” and assume “validate” is 10-20% overhead. In practice, it’s often 50-70% overhead, and it compounds when the governance structure isn’t clear (which it usually isn’t on first AI projects).

4. The Integration and Infrastructure Overhead

The model doesn’t live in isolation. It lives in your product. To make that work, you need to: – Set up infrastructure to serve the model – Build monitoring to track model behavior – Set up logging to understand what the model is doing – Build feedback loops to collect data on what’s actually happening vs. what you predicted – Set up retraining pipelines so the model stays fresh – Build fallback systems in case the model fails

This is often 30-50% of the total project cost, and it’s rarely budgeted separately. It gets rolled into “engineering overhead” or “platform overhead.” Then halfway through the project, the infrastructure work surfaces as a separate project that wasn’t in the original plan.

5. The Expectation and Iteration Tax

You deliver the model. It works. But the business stakeholders have ideas. What if we changed the optimization target? What if we added this constraint? What if we included this data source?

Each of these is reasonable. Each is a couple weeks of work. But if you’ve got three business stakeholders with two ideas each, that’s six iterations that weren’t in the original plan.

This is particularly tricky because it’s not the AI team’s fault. It’s the normal cost of stakeholder collaboration. But it’s not in the budget.

How Projects Actually Fail (It’s Not What You Think)

Projects blow up not because the AI is hard. Projects blow up because one of these five cost categories explodes, and there’s no buffer.

You start with: – $500K for team salaries (data scientists, engineers) – $500K for computing costs – $300K for infrastructure – $700K for contingency

Then you discover in month four that your data quality problem requires reworking 60% of your dataset. That’s $200-300K of work that wasn’t in the “data prep” budget. You absorb it from contingency.

Then you discover in month six that the model architecture you chose doesn’t scale the way you expected. You need to rewrite the serving infrastructure. That’s another $150-200K. Contingency again.

Then your business stakeholder wants to optimize for a different metric because the business priorities shifted. That’s another $100-150K of model work and validation.

By month eight, your contingency is gone, your data team is burned out because they’ve been fighting data issues for six months, your infrastructure team is frustrated because they’re building stuff that’s changing based on model decisions, and you’ve got $500K left and need $800K more to finish.

The project doesn’t fail because AI is hard. It fails because the budgeting framework assumed these categories of work were known and fixed, when they’re actually variable and interdependent.

What Budgeting Actually Needs To Account For

If you’re going to budget for an AI project correctly, you need to plan for:

  1. A discovery phase with explicit budget. Separate from the build phase. If you’re going to learn 40% of what you need to know from building, budget that as learning, not execution. This might be 2-3 months and $300-400K for a medium-sized project. The output is not a production model. It’s clarity on what the production model will require.

  2. Variable data and exploration costs. Budget for data work separately and generously. Assume you’ll need to rework data. Assume you’ll need more data than you think. Assume you’ll iterate on features and labeling. A 40% buffer on data costs is normal, not unusual.

  3. Governance and validation as a distinct work stream, with separate ownership. Not a 10% overhead tax. A separate team, separate timeline, separate budget. If you’ve got a data scientist building a model, you need a governance architect validating it. That’s a different resource, different cost.

  4. Infrastructure and serving as a separate project. The model is not the product. The system that makes the model production-ready is. Budget these separately. Assume infrastructure will be 30-50% of the total project cost if you’re building new.

  5. Explicit contingency for stakeholder iteration. If your business doesn’t know what success looks like, budget for them to figure it out while you’re building. That might be 1-2 iterations that weren’t in the spec. Budget them.

  6. Explicit assumptions about model performance. Don’t budget based on “build a model.” Budget based on “build a model that achieves X accuracy on Y metric while satisfying Z constraints.” The constraints matter more than you think, and they cost real money to validate.

The Math That Actually Works

Here’s what a real budget might look like for a $2M AI project:

  • Discovery phase: $300K, 2-3 months. Goal: understand data, validate problem, identify architecture approach.
  • Build phase: $800K, 4-5 months. Build model, do initial validation, iterate on architecture.
  • Governance and validation: $300K, 2-3 months. Deep validation, bias testing, failure mode analysis, documentation.
  • Infrastructure and serving: $400K, 3-4 months. Build serving infrastructure, monitoring, retraining pipelines, integration.
  • Iteration and contingency: $200K, ongoing. Handle stakeholder changes, unexpected issues, performance improvements.

Notice: the “model” is maybe 40% of the cost. Everything else is making the model real and acceptable.

Notice: governance is explicit, not overhead.

Notice: there’s actual contingency, and it’s for things you expect to happen.

This is a better framework not because it predicts perfectly. It’s better because it accounts for the actual sources of cost in AI projects, and it gives you room to manage them when they happen.

Your next AI project will blow up. Some part of it will cost more than you expected. The question is whether you budgeted for that or whether you pretended it wouldn’t happen.

In Practice: AI in the Enterprise | Day 12: The Hiring Mistake That Sinks AI Programs: You’re Looking for Unicorns

Every large organization I’ve worked with has run the same hiring search: “Looking for an AI/ML Leader with 10+ years of machine learning experience, deep understanding of governance, strong executive presence, ability to translate technical concepts for boards, proven track record scaling data science teams, and experience in regulated industries.”

They’re looking for a unicorn. And it’s sinking their programs.

The person who is genuinely world-class at machine learning research isn’t going to want to spend 40% of their time in governance meetings. The person who is genuinely world-class at organizational design and governance isn’t going to have a publication record in top machine learning conferences. The person who is genuinely world-class at executive communication isn’t the same person who thinks deeply about model architecture.

You don’t need one person to be all three. You need a different organizational structure.

What Companies Actually Need (But Don’t Know How to Ask For)

Let me separate the actual functions that need to exist:

Function 1: Model Builder. Someone who understands how to train, fine-tune, or integrate foundation models. Understands technical tradeoffs, knows how to debug model behavior, can estimate computational requirements and inference costs. This person is a technologist. They might have a PhD in ML. They might have 5 years of production AI experience. They might have started last year and be brilliant. The key trait: they know AI systems can be built and how to actually build them.

Function 2: Program Governance Architect. Someone who understands how AI systems fail in organizational contexts, how to build accountability structures, what risks matter vs. don’t, how to translate technical capabilities into business constraints, how to design approval processes that actually work. This person might not be able to build a model. But they understand organizational design, they’ve worked in regulated industries, they’ve seen programs fail and know why. They’re probably not an ML researcher. They’re probably someone with operating experience and a deep understanding of governance.

Function 3: Executive Translator. Someone who can articulate what’s possible, what’s risky, what’s strategic about AI to boards and CFOs. This person might not understand backpropagation. They might not care about model architecture. But they understand business impact, board dynamics, shareholder communication, how decisions actually get made in the C-suite. They know how to translate “the model drifted” into “our margin impact is X.”

What organizations usually do is try to hire one person for all three. The result is: you hire someone who’s adequate at all three, great at none.

Why The Unicorn Hire Fails

The unicorn hire feels safer than the distributed model because there’s one person who “owns” AI. One P&L. One budget. One accountability line. Your board likes this. Your CFO likes this. It’s clean.

It fails for three reasons.

First: you’ve created a single point of failure. If the person leaves, or underperforms, the entire program slows or stops. You haven’t actually built organizational capability. You’ve hired individual capability. And individuals leave.

Second: you’ve created role confusion. The person is trying to do three jobs. They prioritize the one they like or are good at and deprioritize the others. Usually they deprioritize governance. Then you’re shocked to discover 18 months into your program that you have zero governance structure. This isn’t because they’re bad at governance. It’s because they’re spending 70% of their time building models and 20% of their time talking to executives. Governance gets 10%.

Third: you’ve hired against the market you actually face. The market doesn’t have people who are world-class at model building and organizational design and executive communication. So you hire someone who’s medium at all three. Six months in, you discover they’re not technically deep enough for your hardest problems, or not organizationally sophisticated enough for your actual governance needs, or not credible enough at the executive level. You’ve filled the role without solving any of the underlying problems.

What Actually Works: The Supporting Structure

The organizations that get their AI programs right have a different structure. It’s not one leader. It’s typically three people or roles working in explicit collaboration:

The Chief AI Officer or AI Program Lead is the integration point. They own the P&L, they own the roadmap, they manage the budget. But they don’t try to be the world’s best at everything. They’re typically 40-50% governance architect, 30-40% operator/executor, and 20% translator. They’re comfortable with not being the smartest person in the room on pure ML.

The Technical Leader (whether that’s a VP of Data Science, Head of Applied AI, or Principal Architect) is responsible for the actual AI/ML execution. They own the model builds, the infrastructure, the performance. They’re not trying to be a board-level communicator. They’re not designing organizational structures. They’re focusing on technical excellence and making sure the model systems work.

The Governance / Risk Lead is responsible for designing the structures, processes, and accountability mechanisms. They might sit in the CAIO’s organization, or they might sit in Risk/Compliance/Legal. But they’re explicitly responsible for governance design, not just “making sure we follow the rules.”

And there’s a Chief Executive / CFO / Board Member who is explicitly responsible for asking hard questions, making sure these three are actually collaborating, and understanding what the actual risks and opportunities are.

Notice what this structure has that the unicorn hire doesn’t: explicit accountability for each domain, no single person trying to be great at everything, and organizational incentives for collaboration rather than role confusion.

How to Actually Hire For This (If You Want To Start Now)

If you’re going to build this structure, the hiring conversation is different.

For the Chief AI Officer, you’re not looking for “deep ML expertise.” You’re looking for someone who understands organizational design, has operated in scaled environments, can read a balance sheet and understand how to translate technical decisions into financial impact, and can manage a diverse team. You probably want someone with 10+ years of operating experience. They might have never trained a model. That’s fine.

For the Technical Leader, you’re looking for someone who understands AI/ML systems and can ship them. Deep expertise in your specific domain is nice but not required. You want someone who can articulate technical tradeoffs, who has hands-on building experience (not just theoretical), and who can mentor the team. You probably want 5-8 years minimum, but that’s experience doing something, not background.

For the Governance Lead, you’re looking for someone who understands organizational risk, processes, and compliance. They might come from traditional software governance, or from regulatory work, or from program management. They need to understand what it means for an approval process to actually work. They don’t need to understand backprop.

And you need to explicitly structure their collaboration. Quarterly sync. Shared OKRs. Regular forums where disagreements get surfaced and resolved. Not friendship. Collaboration.

Why This Matters For Your Board

The reason this matters is that it changes how you talk about AI risk and opportunity. With the unicorn model, you’re always one person away from “what happens to our AI program?” With the distributed model, you have institutional capability.

It also changes how you think about scale. One person can manage one program. Multiple people can scale across multiple programs, multiple business units, multiple risk domains.

And it changes how you assess whether you’re actually solving the governance problem. With the unicorn, you can’t tell. With the distributed model, you can see which function is working and which is breaking.

This isn’t organizational theory. It’s practical. The organizations that have figured this out are the ones whose AI programs are actually running, generating value, and not creating constant crises.

In Practice: AI in the Enterprise | Day 11: The Foundation Model Dependency Trap (And Why It Matters More Than You Think)

A CTO I know recently described a conversation with their board: “We’ve built our entire competitive advantage on a foundation model API. What happens when your vendor changes their pricing? What happens when they change their terms? What happens when they decide to compete with us directly?”

The board had no answers. Not because they’re unprepared. But because nobody had modeled the scenario. It felt theoretical, or distant, or like someone else’s problem.

It’s not. It’s the central architecture problem facing enterprises right now, and most companies haven’t thought through the implications.

The Vendor Lock-in Is Structural, Not Accidental

Let’s be precise about what’s happened in the last 18 months. Foundation models have become genuinely useful, genuinely reliable, and genuinely easy to integrate into business systems. A team of three engineers can build a production system that was impossible to build in 2022.

That’s the good news.

The dependency is also structural, not accidental. You’re not using a foundation model as one option among five. You’re building on top of it because:

  1. The capability is native to the model. You can’t replicate this behavior with a smaller model or open-source alternative because the few-shot reasoning, the instruction-following, the broad knowledge cutoff—these are unique to models of this scale.

  2. Your domain adaptations are optimized for that specific model. You’ve tuned prompts for its strengths. You’ve built workflows that depend on its specific output format. You’ve designed your architecture around its API constraints. Moving would require rearchitecting.

  3. Your team has built muscle memory. Your engineers think in terms of tokens. Your product managers understand the API’s quirks. Your operations team has learned how to handle the cost profile. There’s switching cost beyond the code.

When you’ve integrated this deeply, it stops being “we can swap vendors if we need to” and becomes “this vendor is our foundation.”

The question that should terrify you isn’t hypothetical: What’s your actual downside if the terms change?

What “Terms Change” Actually Looks Like

Let me name some scenarios that have already happened:

A foundation model provider raised pricing on their flagship API, and companies that had modeled economics around older versions had to reprice their products or absorb margin compression.

A foundation model provider changed their API rate limits, and companies hit unexpected bottlenecks in their production systems.

A foundation model provider changed their token-counting methodology, and billing shifted for systems that hadn’t anticipated the change.

A foundation model provider started offering cheaper models with similar capabilities, and companies that had built competitive advantages on the assumption of exclusive access to high-quality reasoning were suddenly competing on cost instead.

A foundation model provider changed their pricing on advanced context window features, and companies had to rebuild systems they’d optimized for that specific cost structure.

The pattern is consistent: these aren’t conspiracy theories or paranoia. These are normal business decisions from vendors optimizing for revenue and managing demand. And every time they happen, they force expensive recalibration on dependent companies.

But the scenario that should worry you more is the one that hasn’t happened yet: what if your foundation model provider decides to build a competing product?

The Unspoken Competitive Exposure

Here’s the asymmetry that matters: You’re optimizing your systems on their infrastructure. Your customer relationships, your domain expertise, your operational knowledge—these all live on top of their API.

They see your usage patterns. They see what you’re building. They understand your market. They have the infrastructure. They have the customer relationships.

If they decide your market is valuable, they can build a competing product faster, cheaper, and with built-in advantages (integration with their foundation model, direct access to their API, no external dependencies, direct customer contact for switching). You’ve essentially market-tested the opportunity for them.

This isn’t hypothetical either. It’s happened in cloud computing (major cloud providers started with foundational services but moved up into competitive applications). It’s happened in search (dominant platforms built ad networks, then started competing with the services that were built on top of search traffic). It’s the normal trajectory of platform companies.

The difference here is that the switching costs work in the vendor’s favor. If you’ve built an integration layer for one vendor’s API, moving to another foundation model provider isn’t just an engineering project. It’s a rebuild. They know this. You should price this risk into your strategy.

What Most Companies Are Actually Doing (And Why It’s Insufficient)

I’ve seen three patterns in how enterprises are thinking about this.

The Hedging Approach: “We’ll use multiple foundation models.” This reduces the risk that any single vendor’s downtime breaks your system. It doesn’t reduce dependency. You’re now dependent on multiple vendors, each of whom has independent incentive to compete with you or change their terms. It also increases operational complexity and cost. You’re not building redundancy; you’re building technical debt.

The Wait-and-See Approach: “We’re starting with one foundation model, and we’ll re-evaluate in a year.” This is rational caution. It’s also how you end up with deep integration by the time you realize you need an alternative. Year-one projects become standard. Dependencies calcify. By the time you re-evaluate, your systems have dependencies you forgot you had.

The Open-Source Alternative Approach: “We’ll train or fine-tune our own models.” This is intellectually satisfying and sometimes necessary. But most enterprises dramatically underestimate the cost. Training a capable foundation model requires infrastructure, talent, and capital that’s only cost-effective if you can amortize it across many applications or many customers. For a single product or department, it’s usually more expensive than staying dependent.

None of these approaches actually solves the dependency problem. They just mask it differently.

What Changes When You Design for Portability

The companies that are handling this well have made a structural decision: they’re building abstraction layers.

Instead of writing code that assumes a specific foundation model’s API response format, they write code against an abstraction that any foundation model interface could satisfy. It’s more code. It’s slightly less performant. It’s definitely more complex in the short term.

But it has one critical property: if the economics of your primary model change, or if the capabilities diverge, or if a better alternative emerges, you can actually migrate. It’s not effortless. But it’s possible without rebuilding your product.

This requires:

  1. An abstraction layer at the application boundary that decouples your code from specific model APIs.

  2. A cost model that includes the switching overhead in your unit economics. If you can only switch at 40% cost premium, that’s your real cost baseline, not the current API pricing.

  3. A technical strategy that explicitly plans for model rotation. Not “if” but “when.” Assume you’ll need to switch models at least once over the product’s lifetime.

  4. Monitoring around model performance that’s independent of vendor claims. You need to know when a model’s behavior changes, not when the vendor announces it.

This isn’t about paranoia. It’s about building businesses that can adapt when the infrastructure under them changes. Platform dependencies are real, and they’re not unique to foundation models. But foundation models are new enough that most enterprises haven’t yet built the organizational practices to manage them.

The Conversation You Should Be Having Now

The question isn’t “should we be dependent on a foundation model?” You should be. They’re better at many tasks than anything you can build or buy.

The question is: “What’s our dependency management strategy?”

That’s a conversation between your CTO (who owns the architecture), your CFO (who owns the total cost), your product head (who owns customer impact of changes), and your board (who owns risk). It’s not primarily a technical conversation. It’s a strategy conversation.

The companies that will compound value from AI over the next five years won’t be the ones that built the fastest integration in 2024. They’ll be the ones that built it in a way they can still change in 2026.

In Practice: AI in the Enterprise | Day 10: Why “Monitoring” Your AI Means Something Different Than Monitoring Your Databases

There’s a moment in every enterprise AI deployment when someone asks: “So we’ll monitor the model, right?”

And then they spend the next meeting explaining what they mean by that.

It’s one of those words that sounds simple but means completely different things depending on who’s using it. When a data engineer says they’re “monitoring” a database, they mean something specific: is it up? Is it responding? Are queries returning in acceptable time? Is storage growing unexpectedly?

When a data scientist says they’re “monitoring” a model, they usually mean: is the model’s performance degrading? Is the data distribution changing? Are we hitting the thresholds we set?

These are valuable questions. But they’re also not the questions that usually matter most in enterprise AI.

Here’s what I’ve learned: the guardrails that fail first in production are not the ones built on model monitoring. They’re the ones built on monitoring the wrong signals.

Most enterprises build their AI observability around technical metrics. Model performance, drift detection, data quality. These are important. But they’re also the ones that usually don’t surface the problems that actually matter.

Let me give you a concrete example of what I mean.

A financial services company deployed a system to recommend investment allocations. The model was good. It predicted portfolio performance accurately on historical data. In production, they monitored drift: is the model’s performance degrading? They monitored feature distributions: is the data it’s seeing different from what it trained on?

For three months, everything looked fine. No drift. Data distributions stable. The model was performing exactly as expected.

But during month four, a compliance officer pulled a sample of the system’s recommendations and noticed something. The allocations were technically sound, but they were increasingly concentrated in a narrow set of holdings. The model wasn’t breaking. But it was drifting behaviorally—making increasingly concentrated recommendations without the organization realizing it.

They only caught it because someone looked at actual recommendations. Not because the monitoring systems flagged anything.

Here’s what actually needs to be monitored in production AI systems:

First, signal integrity. Is the model receiving the data it expects to receive? Not just: is the data schema correct? But: is it coming from the same sources? Are the sources themselves changing? I know of a system that stopped working not because the model broke, but because the upstream data pipeline was silently modified and suddenly all the features were null. The model kept running. It kept producing predictions. Nobody noticed for weeks because they weren’t monitoring whether the input signals were actually present.

Second, decision consistency. When the model makes decisions, are they consistent with what you’d expect? This is qualitative. You can’t automate it. But you need someone periodically looking at actual decisions the system is making and asking: does this look right? I saw a recommendation system that was technically working fine but had learned to make increasingly extreme recommendations to a small segment of users. The metrics didn’t catch it. A human looking at recommendations caught it immediately.

Third, outcome correlation. You built a system to predict X and optimize for Y. In production, are the outcomes you’re getting consistent with the optimization? Or has the system learned a pattern that makes the metric go up but doesn’t produce the desired outcome? I’ve seen models that minimized their loss function beautifully but maximized customer churn because they learned a correlation that was mathematically true but operationally wrong.

Fourth, distribution shift that matters. Most systems monitor whether the input distribution has changed. But not all distribution shifts matter equally. A shift in the age distribution of your customer base might matter. A shift in the geographic distribution might not. You need to monitor distribution shifts relative to what you’re trying to predict, not just whether anything changed.

Fifth, the rare cases. Your model works great on the common case and terribly on the rare case. In production, are you tracking how the model behaves on the cases that are actually rare and important? Or are you only looking at average performance across everything? A lot of systems fail silently on the 2% of cases that actually matter most.

The organizations that are ahead on AI observability aren’t the ones with the most sophisticated monitoring systems. They’re the ones that are clear about what signals actually matter.

And those signals usually aren’t in your standard model monitoring dashboard.

Here’s what I’d recommend: build your monitoring in layers.

Layer 1: Infrastructure monitoring. Is the system up and running? Is it fast? This is table stakes. Use standard database and application monitoring.

Layer 2: Technical drift monitoring. Is the model’s performance degrading? Is the input data distribution changing? These are important. Use standard MLOps tools.

Layer 3: Decision monitoring. What is the system actually deciding? Are the decisions consistent with what we expect? Is the decision distribution drifting? This usually requires custom instrumentation.

Layer 4: Outcome monitoring. Are the outcomes we’re getting from the system aligned with what we wanted? This requires tracking the actual business outcomes, not just the system’s metrics. This is often the hardest layer to build, but it’s the most important.

Most enterprises have Layer 1 and Layer 2 down. They’re weak or absent on Layer 3 and 4.

That’s where the guardrails fail.

The reason is simple: Layer 1 and 2 are about the system. Layer 3 and 4 are about what the system actually does in the world. And those are different problems with different solutions.

So when someone says “we’re monitoring the model,” ask: which layer are you monitoring? If the answer is “just technical drift,” you’re missing where the real risk lives.

Build guardrails around signals that matter. Track actual decisions. Watch for outcomes that diverge from intent. Layer those on top of your standard monitoring.

The systems that fail quietly aren’t the ones where the model broke. They’re the ones where nobody was watching the right signals.

In Practice: AI in the Enterprise | Day 9: The One Document Your Legal Team Should Demand Before Any AI Goes Live

I was in a meeting with a general counsel last month. The company was deploying a large language model as an internal knowledge assistant. It was a reasonable project. The team had thought through data security. They’d built audit trails. They’d set usage policies.

The GC asked a question that stopped the room: “What happens if the model hallucinates and gives someone bad advice they act on?”

Not a weird question. A straightforward question. And it turned out nobody had a clear answer.

The conversation that followed was instructive. The data science team said: well, it’s a language model, hallucinations are a known limitation. The product team said: we have a disclaimer. The legal team said: disclaimers don’t protect us from liability.

And that’s where it usually breaks down.

Many organizations understand that AI systems can fail. What many are still developing their understanding of is the structure of liability when they do.

Here’s what I think you need before any AI system goes live in your organization:

A hallucination liability assessment.

This isn’t model documentation. This isn’t a risk register. This is a specific document that your legal team writes, not your data science team, that answers one question: if this system confidently says something that’s false, and someone acts on it, what is our legal exposure?

That’s the starting point. Because hallucination—when a language model generates confident falsehoods—is a specific kind of failure mode that most organizations haven’t built legal thinking around yet.

For decades, organizations have built legal frameworks for system failures: databases go down, processes fail, data gets lost. These are understood. There are precedents. There are insurance products.

But a system that confidently tells someone something false is a newer problem. And the legal structures around it are still settling.

Here’s why it matters:

If your system is purely informational (you’re using it to summarize documents internally), the risk profile is one thing. Someone reads the summary, notices it’s wrong, corrects it. Inconvenient. Not necessarily a legal problem.

But if your system is integrated into a workflow where a human is supposed to verify the output, but the human reasonably relies on the system’s apparent confidence, you have a different risk structure.

And if your system is trained on your proprietary knowledge and someone acts on its hallucinatory output in a way that damages a customer relationship, a business deal, or a compliance outcome, you have another risk structure entirely.

These aren’t hypothetical. I’ve watched organizations have exactly these moments. “The model told our sales team that we offered a service we don’t actually provide.” “The system generated a draft contract clause that contradicted our terms but the team didn’t catch it.” “The model synthesized customer data in a way that suggested we had capabilities we don’t have.”

None of these broke the system. All of them created liability.

So before you deploy, you need legal to answer:

First, context: What decisions or actions does this system enable or inform? What’s the chain from system output to human decision to real-world consequence?

Second, confidence: When this system is wrong, is it obviously wrong? Or does it confidently present falsehoods? (This matters because overconfident falsehoods create liability that uncertain outputs don’t.)

Third, verification: In the workflow, is there a verification step? And critically, is that step designed to catch hallucinations, or just to audit that the system ran?

Fourth, scope: Who can see the output? Is it constrained (internal team) or broad (customers, partners)? Scope multiplies liability.

Fifth, exposure: If the system hallucinates and someone acts on it, what’s the worst-case consequence? Client relationship? Compliance violation? Wrong regulatory filing? Dollar impact? This determines what you need to mitigate.

Once you’ve answered these, you’re in a position to actually assess whether you should deploy this system as-is, or whether you need to change something about how it’s used.

Maybe you need human verification at a specific step. Maybe you need to constrain who can access the output. Maybe you need specific disclaimers that are actually meaningful. Maybe you need to train people on the limitations. Maybe you need to not deploy it the way you initially planned.

Most organizations don’t do this. They get to deployment and they ask: does it work? And if it works, they deploy it. Then they’re surprised when the liability question comes up in month four because someone acted on something the system hallucinated.

Legal teams are increasingly asking for this, but they’re often asking late. The system is built. The deployment is planned. Now they’re trying to add guardrails to something that was designed without legal thinking.

That’s expensive and it’s slow. Better to do it upfront.

Here’s what I recommend: before you move a large language model (or any AI system with significant hallucination risk) from pilot to production, have your legal team write a one-page assessment. Not a binder of documentation. Not a risk matrix. One page. What’s the hallucination exposure? What are the paths from model output to liability? What mitigations do we need?

Then, once you have that clarity, the engineering team can decide: do we change the workflow? Do we add verification? Do we constrain access? Do we add specific training? Do we add disclaimers that actually mean something?

The companies that are handling this well aren’t the ones with perfect models. They’re the ones that are clear about the risk, and clear about what they’ve done to manage it.

And that clarity starts with a document that your legal team owns, not your data science team.

Get that document written before you go live. It will change how you deploy. And it might prevent you from being the organization that learns the liability lesson the hard way.

In Practice: AI in the Enterprise | Day 8: The Bank That Realized (Too Late) Their AI System Was Making $2M Mistakes in Production

It was a Wednesday afternoon in late October when someone noticed something odd.

A risk officer was reviewing a batch of decisions made by a lending recommendation system. The system had been live for eight months. It was working. Models had been validated. Metrics looked good. But in this particular batch, something didn’t add up.

There was a pattern of approvals that shouldn’t have been approved.

The system had learned to optimize for one thing (approval velocity), and in doing so it had learned to weight certain borrower profiles in a way that made internal sense but created external risk. It wasn’t broken. It wasn’t defrauding anyone. It was just… quietly making decisions that looked right by the metrics the organization was monitoring, but created risk the organization hadn’t accounted for.

When they dug in, they found that over eight months, this system had recommended approval decisions on credits that had characteristics which, under scrutiny, suggested higher default probability than historical norms. The organization had essentially been taking on tail risk that nobody had measured because nobody was measuring the right thing.

The cost, when they finally calculated it, was in the millions. Not because the model was wrong. But because the operational reality of how the system was being used diverged from how it was supposed to be used.

This is not a story about a specific bank. But it’s a pattern I’ve seen map across different domains, different risk types, different sizes of organizations. And it illustrates something fundamental that I think enterprise leaders chronically underestimate: operational risk in AI systems is different from model risk.

Model risk is the question: does the model predict accurately? You can measure that. You can validate it. You can test it on hold-out data.

Operational risk is the question: what actually happens when humans and machines make decisions together in the real world, over time, at scale?

It’s much harder to measure. And it’s where the real money goes.

Here’s how it usually happens:

The system is deployed. It works. Engineers and data scientists are satisfied because the metrics are holding. The business is satisfied because the system is processing decisions faster. Everyone feels good.

Then, slowly, a gap opens up between how the system is being used and how it was designed to be used.

Maybe a human decision-maker starts trusting it too much and stops applying their judgment. Maybe they start using it in cases it wasn’t designed for. Maybe the process changes and the system isn’t updated. Maybe the data it’s learning from drifts because the business has shifted and nobody told the model.

These aren’t failures. They’re just reality. Real humans, real workflows, real organizations. Things change. People adapt. Systems… sometimes don’t.

And in the gap between how the system was designed and how it’s actually being used, risk accumulates. Silently.

The lending scenario is clean because the risk is quantifiable. But the pattern is broader.

A claims adjudication system that learned to approve claims faster starts approving claims that shouldn’t be approved. A hiring system that optimizes for filling positions quickly learns to screen out categories of applicants that were never the intent. A fraud detection system that minimizes false positives starts letting real fraud through. A customer retention system that minimizes churn learns to make offers that are economically senseless.

None of these happen because the model broke. They happen because the operational reality diverged from the model’s intent.

So how do you prevent it?

The honest answer is: you can’t, completely. Humans are adaptive. Systems are rigid. Gaps will open.

But you can reduce it dramatically by doing something most organizations don’t do: operationalize decision monitoring.

Not model monitoring. Decision monitoring.

Model monitoring is: is the model drifting? Is the data quality declining? These are good questions. Most enterprises do some version of this.

Decision monitoring is different: what decisions is the system actually making? Are they consistent with what we expected? When the system and humans disagree, what’s the pattern? What are we learning about how this system is actually being used?

This requires infrastructure that most organizations don’t have. You need to log not just the model’s prediction, but the human’s decision. You need to understand the gap between them. You need to understand how that gap is changing over time.

Then you need someone—an actual human with authority—reviewing this regularly. Not algorithmically. Manually. Reading through cases. Asking: does this feel right? Where should I worry?

The lending organization I referenced earlier, once they figured out what had happened, implemented this. They assigned a risk officer to review 50 random decisions per week from the system. Just reading them. Just asking: do these approvals make sense? Do they align with our risk appetite?

That single investment—one person’s time—would have caught the drift before it became millions in tail risk.

Most organizations won’t do this. It feels expensive. It feels like overhead. The system is working. The metrics are good. Why do you need a human reading cases?

Because operational risk isn’t in the metrics. It’s in the gap between what you’re measuring and what’s actually happening.

Here’s the hard part: you probably can’t see that gap from the metrics. You have to see it by looking at actual decisions, actual outcomes, actual patterns in how the system is behaving in the world.

So if you deploy an AI system that makes real decisions—lending, hiring, fraud, claims, pricing—you need someone whose job is to catch the operational risk. Not to prevent the model from drifting. To catch the moment when the organization’s use of the system starts diverging from the system’s design.

It’s boring work. It won’t show up in quarterly metrics. It’s also how you avoid being the organization that discovers, in month eight, that you’ve been taking risks you didn’t know about.

The cost of operational risk isn’t usually in one bad decision. It’s in the pattern of decisions that are individually defensible but collectively add up to exposure you didn’t want.

Pay attention to that gap. Assign someone to watch it. It’s cheaper than learning the lesson the hard way.

In Practice: AI in the Enterprise | Day 7: You’re Measuring AI Success Wrong, and Here’s the Cost

Every data science team has a graveyard.

You won’t see it in the org chart or the budget forecasts. It’s the projects that worked—on paper. They hit their targets. The models converged. The metrics looked clean. Then they got four months into production and somehow, without anyone quite understanding when it happened, they became optional.

A recommendation engine that no one acts on anymore. A churn prediction model that technically has good precision but the business stops using it because it “doesn’t feel right.” A customer segmentation model that’s accurate by every measure except the one that actually matters: whether the business can do anything useful with the segments.

If you ask the data science team what went wrong, they’ll tell you the truth: nothing went wrong. The metrics didn’t move. The model didn’t degrade. Everything performed as expected.

And that’s the problem.

The wrong metrics don’t fail. They persist. They sit there, humming along, proving they work, while the business quietly stops trusting them.

Most enterprise teams measure AI success through a lens designed for a different problem: Can we build a model that predicts X with accuracy Y?

That’s a data science problem. It’s not an enterprise problem.

An enterprise problem looks like: Can we deploy a system that helps a human make better decisions, at a cost we’re willing to pay, in a way that creates actual value?

These are not the same question. And the gap between them is where value goes to die.

Here’s what happens. Your team builds a credit risk model. They measure it on hold-out test data. Precision, recall, AUC—all of it beautiful. 0.92 AUC. You can draw the ROC curve and it’s smooth and clean.

Then it goes live. And in month two, a credit analyst pulls you aside and says: “The model is right more often than I am. But I’m the one who has to justify the reject to the customer. When the model says no and I say no, I sound like a robot reading numbers. But I can explain why in a way they trust.”

And suddenly that 0.92 AUC doesn’t matter. What matters is whether the credit analyst will use the model or work around it.

Most organizations don’t measure that. That’s not a metric. That’s an adoption problem. And adoption problems are usually blamed on the humans, not the model.

“If they’d just trust the model more,” you hear. “If they understood how accurate it is.”

But the humans understand the metrics perfectly. They’re just answering a different question: does this reduce my risk while I do my job?

The real cost of measuring success the wrong way shows up later. You’ve deployed 47 models. By your metrics, 44 of them work. By the organization’s metrics—are they actually being used in production decisions, are they creating value, do people trust them—maybe 12 work.

You’ve spent millions on infrastructure, data science teams, model governance frameworks. You’ve built a technically sophisticated machine that delivers technically perfect predictions nobody cares about.

This is why so many enterprise AI programs feel expensive and slow and produce results that never quite justify the investment.

It’s not that the organizations are incompetent. It’s that they’re measuring success at the wrong level.

You need at least three measurement frameworks running in parallel:

The first is technical. Does the model predict accurately on hold-out data? Can it handle the data it receives in production? Does it drift? These are table stakes. But they’re not sufficient.

The second is behavioral. When the model makes a recommendation or a decision, what does the human do? Do they follow it? Do they override it? Do they use it as input to their own judgment, or do they ignore it? How often do they come back and say “you were right” vs. “that was wrong”? If your model is predicting well but humans override it most of the time, you don’t have a model problem. You have a deployment problem. But you only see it if you’re measuring it.

The third is economic. Does using this model reduce costs, increase revenue, or improve customer outcomes by an amount greater than what we spent building it? You’d be shocked how many enterprise teams don’t know the answer to this question. They have a model. It works. But has it actually paid for itself?

The cost of getting this wrong is two-fold. First, you waste money on systems that don’t deliver value. Second, you damage the credibility of the whole enterprise AI program. You build a reputation for expensive, sophisticated systems that don’t work. And “work” in the minds of decision-makers means: does it help me do my job better?

The organizations measuring AI success effectively aren’t the ones with the best data scientists. They’re the ones that are clear about what “success” means before they start. They’re clear about it in a way that’s measurable, observable, and tied to something the business cares about.

A head of AI I know runs a quarterly check on every model in production. It’s simple. She asks: Is anyone using this? Are they using it the way we expected? Is it creating value? If any of those answers is no, they either fix it or they retire it.

No heroes. No “it’s technically correct so we must be doing something wrong.” Just: does this work in practice?

That discipline—measuring success in three dimensions, and being willing to retire systems that work on metrics but not in reality—is what separates programs that produce value from programs that produce impressive dashboards.

Your next AI project? Decide what success looks like before you start. Make sure at least one dimension of it is something humans have to care about. Then measure it, all the way through.

The metrics will take care of themselves.