In Practice: AI in the Enterprise | Day 26: What “Responsible AI” Actually Means in Production (Spoiler: It’s Boring and Unglamorous)

In the conference talk, “responsible AI” sounds like something you’d frame on your wall. A commitment. A value. A north star.

In production, it looks like this: A spreadsheet. A checklist that nobody finds interesting. Someone logging into a monitoring dashboard at 3 a.m. because an anomaly flagged and they need to know whether the model just broke or whether the data distribution shifted.

That’s not the story we tell about responsible AI.

The story we tell is about bias detection and mitigation frameworks. About fairness metrics and algorithmic transparency. About governance structures and ethics review boards. These are all real. These are all important. But here’s what I’ve observed in organizations that are actually running AI systems in production: The part that actually makes a difference isn’t any of those things.

It’s boring operational discipline.

I was talking to an engineering leader at a financial services organization who’d implemented what looked like an exemplary responsible AI program. They had fairness audits. They had bias mitigation protocols. They had a governance review process that was genuinely rigorous. And yet, they kept discovering problems.

Not bias problems. Not fairness problems. They discovered issues because their monitoring was good enough to catch something changing, and then they had to figure out what changed and why.

The issue wasn’t their governance framework. The issue was that they’d built a system that could see what was happening in production. That visibility—that’s what made them responsible.

Here’s the distinction that matters: Responsible AI in theory is about making good decisions about whether to deploy a model. Responsible AI in practice is about having enough visibility into what the deployed model is actually doing that you can react to it.

The shift is from governance (approval before launch) to observability (understanding after launch).

Think about what this means. Many organizations run AI systems the same way they’d run software ten years ago. They test. They deploy. They check some metrics. If the metrics look good, they move on. If something breaks, they investigate.

Responsible AI in production requires a different operational model. It requires continuous monitoring of:

Whether the model’s outputs are changing in ways you’d expect or not expect. Not just aggregate accuracy—that can hide problems. Accuracy across different input categories. Accuracy across different user segments. Accuracy for different decision types. The goal isn’t a single number. The goal is a detailed map of where the model works and where it doesn’t.

Whether the inputs the model is receiving are changing. You train a model on data from 2024. By 2025, the input distribution might have shifted in ways that make the model less reliable for certain cases. If you don’t know the input distribution is changing, you won’t know the model is becoming risky.

Whether the model’s predictions are correlated with things they shouldn’t be correlated with. This is where bias detection actually matters operationally. Not as a predeployment audit, but as a continuous check. If the model’s decisions are clustering differently across demographic segments in production than they did in testing, you need to know.

Whether edge cases are being handled the way you expect. Models often fail quietly. Not a crash. Not an error. Just a decision that’s worse than it should be. You need to know when that’s happening and where.

Many organizations lack this level of visibility. They have aggregate metrics. They have dashboards that show “accuracy is at 87%.” But they don’t have models of what each decision is actually doing and whether it’s doing what you’d expect it to do.

Building that visibility is expensive. It requires:

Instrumentation. You need to log not just the model’s predictions, but the inputs, the context, the ground truth (when you get it), the business outcome. All of it. This adds complexity to your serving layer.

Analysis infrastructure. You need systems that can look at all that data and identify patterns and anomalies. Not just “accuracy went down” but “accuracy went down specifically for these inputs and these decision types.”

Alerting and response. When something looks off, you need to know. And you need to know quickly. That means someone is paying attention. That means you have a protocol for investigating and responding.

Integration with your data pipeline. The model isn’t separate from your data. It’s downstream of collection, cleaning, transformation. If something breaks upstream, the model will reflect it. Your observability needs to trace backward through that pipeline.

This is the operational discipline that actually makes AI systems responsible. Not more governance. More visibility.

The honest version of “responsible AI” in production looks like this: We’ve deployed a model that’s going to make decisions that affect real people and real business outcomes. We don’t fully understand how it will behave in all situations. So we’re going to watch it very carefully. We’re going to instrument it in ways that let us see what it’s actually doing. We’re going to set up alarms for when behavior deviates from what we expect. We’re going to log everything so that if a decision deviates from expectations, we can understand why. And we’re going to have a team whose job it is to monitor these signals and respond.

That’s not governance theater. That’s not a checkbox on a compliance framework. That’s not something you can outsource to an ethics board.

It’s unglamorous. It’s expensive. It requires ongoing investment. It requires people.

But it’s also the only thing that actually works.

The organizations doing this well understand that deploying an AI system is the beginning of the work, not the end. The deployment is when you move from theory to reality. That’s when you find out whether your fairness metrics actually predicted real-world behavior. That’s when you discover whether your bias mitigation actually mitigated the specific biases that matter in your context. That’s when you figure out the edge cases that your testing didn’t catch.

And the only way to handle that transition is with visibility. Not oversight. Visibility.

Oversight is about saying “no” before launch. Visibility is about saying “here’s what’s actually happening” after launch. Both matter. But in production, visibility is what keeps you responsible.

The uncomfortable truth about responsible AI is that it’s not something you achieve once. It’s something you maintain continuously. It’s an operational discipline, not a destination. You don’t finish responsible AI. You keep running it.

And that means building systems that let you see what you’ve built, so you can adjust it when it doesn’t work the way you expected.

That spreadsheet I mentioned at the beginning? The one that doesn’t sound impressive?

That’s where responsible AI actually lives.

In Practice: AI in the Enterprise | Day 25: How to Talk to Your Foundation Model Vendor About Risk (When They’re Not Used to Being Asked)

“We haven’t really thought about that,” the vendor says. It’s the third question in a row where that’s their answer.

You’re asking about their model’s training data composition. They give you a vague answer about “internet-scale data.” You ask how they audit for biases in specific demographic segments. They mention some internal metrics but can’t share details. You ask what happens if their model’s outputs degrade over time—how do you know, and what’s their accountability for supporting you through that?

Blank stare.

This is what happens when enterprise risk management meets the foundation model industry. These are questions every organization should ask any model provider—including their own internal platforms. No provider is exempt from this scrutiny.

I watched this unfold at an organization that was evaluating foundation models for a mission-critical use case. They were comparing three vendors. All three had impressive benchmark scores. Two had larger market share. All three struggled when asked basic governance questions.

Vendor A couldn’t explain how they’d support model evaluation in a regulated industry. Vendor B had no framework for understanding when their model was performing poorly for a specific type of decision. Vendor C didn’t have clear terms around what happens if the model causes harm in production—insurance? Liability caps? Shared responsibility?

None of them were bad actors. Many vendor relationships haven’t yet matured to include these conversations.

Here’s what makes this problem acute: You don’t have a choice of who builds the foundation model. But you have full responsibility for how it’s used. Your organization is accountable. Your model’s outputs are driving decisions. Your brand is attached to the outcomes. If the model fails, you’re the one explaining it to regulators, customers, and your board.

But you can’t control whether the vendor invests in understanding failure modes, bias auditing, degradation monitoring, or liability frameworks. You can only choose whether to use their model or not.

So how do you make that choice?

Start by understanding that the vendor’s relationship to risk management is different from yours. They’re thinking about model performance on benchmarks and market adoption. You’re thinking about liability and business continuity. These aren’t opposed—but they’re not aligned either. A model that performs well on benchmarks might have gaps in exactly the areas that matter to your use case.

When you’re evaluating vendors, the technical benchmarks are table stakes. But they should be the end of your technical evaluation, not the beginning. The real questions are about robustness and transparency.

Ask about transparency first. Can the vendor explain what data the model was trained on? In detail, not vaguely. What’s the composition? What’s included and what’s excluded? When was it trained? What happened since? Have there been updates? If so, what changed? This is fundamental. You can’t assess model risk if you don’t know what it was trained on.

Ask about bias auditing. How do they audit for performance gaps across demographic segments? How do they know if the model performs differently for different populations? Do they have data on this? Can they share it? What’s their protocol for identifying and addressing bias? Not “do they think bias is important”—everyone thinks bias is important. “What’s their operational framework for detecting and managing it?” Many won’t have a clear answer. That tells you something.

Ask about degradation. How does the vendor monitor for model drift or degradation? What’s their detection methodology? How quickly would they notice if the model’s outputs started getting worse? What’s their support model if that happens? Some will say they monitor internally but can’t share details for competitive reasons. Others will say they don’t monitor customer deployments at all—that’s the customer’s responsibility. Understand where the boundary is.

Ask about accountability. If the model causes harm in production, what’s the vendor’s responsibility? Is there insurance? Liability caps? Shared accountability? This is uncomfortable to ask because it feels adversarial. But it’s essential. If the vendor hasn’t thought about it, that’s important information. If they have clear answers, that tells you they’ve dealt with enterprises who needed it.

Ask about customization and fine-tuning. If you need to fine-tune the model on your own data, what’s their support? What’s the process? What happens to your data? How do they ensure fine-tuned models maintain their governance properties? Many vendors have weak answers here because they’re optimized for off-the-shelf use, not enterprise customization.

Ask about failure modes specific to your use case. You know better than anyone where this model could fail in your business. Describe it. Ask the vendor if they’ve seen similar issues. What would they recommend? How would they support you if that failure occurred? Some vendors will have thought through this. Others will give you generic reassurance.

Here’s the hard part: The vendor’s answers to these questions are often uncertain or incomplete. They might not have great answers. That’s not a disqualifier. But it’s data. If a vendor has thought deeply about these problems and has frameworks for managing them, that’s worth something. If a vendor hasn’t thought about them at all, that’s also worth knowing. You’re not looking for perfect answers. You’re looking for evidence that they’ve dealt with enterprise risk management.

After you’ve asked these questions, the real conversation starts. Most vendors, when pressed, want to get better at this. They’ve built good models but haven’t had to integrate them into enterprise governance structures. They’ll often work with you on transparency, monitoring, and accountability if you ask clearly what you need.

The organizations doing this well have made a deliberate choice: We’re going to use someone else’s model, but we’re not going to give them our governance responsibility. We’ll understand what we’re inheriting. We’ll set clear expectations. We’ll require transparency and accountability. We’ll integrate their model into our risk management framework, not theirs.

That conversation looks different from a typical vendor evaluation. It’s not about price or feature parity. It’s about understanding whether this vendor can be a partner in managing the risks their model introduces.

Many vendor relationships are maturing into these conversations. And the ones who engage deeply tend to stay. Because when you’re selecting a foundation model, you’re not just selecting a model. You’re selecting a risk partner. And that’s a different conversation entirely.

In Practice: AI in the Enterprise | Day 24: The IP Risk You May Be Building Into Your AI System

Your organization is probably training a model on data you don’t fully own, and you haven’t thought about what that means.

This isn’t theoretical. It’s becoming the reason some organizations can’t deploy models they spent months building, why legal is suddenly involved in product decisions, and why audits are unraveling.

Let me separate myth from reality here, because the legal landscape around training data is messy enough that most organizations are operating on assumptions.

Myth: “If it’s publicly available, we can use it for training.”

Reality: Publicly available is not the same as publicly licensed for machine learning. A dataset scraped from websites is not the same as a dataset licensed under a creative commons agreement. A corpus of text published on the internet is not automatically yours to train a foundation model on. Copyright holders have claims on their work even if it’s freely accessible. Some have already started suing. Others haven’t yet, but they have the legal standing to do so.

The practical implication: You can train a model on public data. You can’t assume you own the right to deploy it without understanding where that data came from and what licenses or claims might attach to it.

Myth: “We’ll just use public datasets that are already curated for ML.”

Reality: Curated datasets often come with documentation about where data came from, but not always with clear licensing for derivative works. And even when they do, that license might constrain how you can use the resulting model. Some licenses allow training but not commercial deployment. Some allow training but require you to make your model public. Some are clearer than others. Most organizations don’t read the license until something breaks.

The practical implication: Even “safe” datasets might not be safe for your specific use case. You need to actually read the licensing terms and understand what they mean for your deployment model.

Myth: “Synthetic data solves the problem.”

Reality: Synthetic data solves one problem—it’s generated from scratch, so you own it. But many approaches to generating synthetic data involve using real data as a seed. If you generated synthetic data by fine-tuning a model trained on licensed data, you’ve inherited the licensing problem. And if you’re using synthetic data that was generated from real data, there are legitimate questions about whether derivative rights claims apply.

The practical implication: Synthetic data is useful, but it’s not an automatic legal escape hatch. Understand how your synthetic data was generated and what licensing applies to the sources.

Myth: “Fair use lets us use any training data we want.”

Reality: Fair use is a legal defense, not a permission. It’s evaluated case-by-case on four factors: the purpose of use, the nature of the copyrighted work, the amount and substantiality of use, and the market impact. Training a commercial AI model is not obviously fair use. The amount and substantiality of use is massive—you’re using the entire work. The market impact could be significant—you might be competing with the original creator. Courts might find fair use applies. Or they might not. That’s not a risk assessment framework; that’s hoping you’re right.

The practical implication: Don’t assume fair use covers your training data. Treat it as a possible defense if you’re sued, not as a license to use whatever data you want.

Myth: “Small amounts of copyrighted data in a large training set are negligible.”

Reality: There’s no legal threshold where “small amount of copyrighted material” becomes acceptable. The question isn’t “how much?” It’s “did you use copyrighted material without permission?” Even a small amount matters legally. Whether it results in actual damages depends on other factors, but the infringement itself is the legal exposure.

The practical implication: You can’t solve attribution problems by dilution. Using a tiny amount of someone’s work without permission isn’t better than using a large amount; it’s still using it without permission.

Now, here’s what makes this complicated: The foundation models you’re fine-tuning or using were trained on data whose provenance is unclear. Major foundation model providers have trained on internet-scale data. They’ve faced legal scrutiny over it. When you fine-tune their model on additional data, you layer that risk on top of new risk.

What should you actually do?

First: Know what data you’re training on. Make a list. Document the source. If you downloaded it, document where. If someone shared it, document that. If you scraped it, document that. Get metadata about licensing and ownership.

Second: For each source, verify the rights. Can you legally use this data for commercial AI training? Not “is it publicly accessible?” but “do we have the rights to train a model on it?” This means reading license agreements. This means understanding copyright claims. This means acknowledging uncertainty when licensing is unclear.

Third: Keep records of your due diligence. If you later discover you used copyrighted material, the fact that you tried to verify rights matters. Willful infringement is worse than inadvertent infringement. Document that you made a good-faith effort to understand what you were doing.

Fourth: For models going into production, have a legal perspective on training data provenance. Not after the fact—before deployment. This is a governance gate, like model validation or safety review. Does legal sign off on the training data used for this model? If not, what’s the risk assessment?

Fifth: For foundation models, understand what you’re inheriting. The model you’re fine-tuning has training data you don’t fully own. When you fine-tune it, you’re creating derivative works of derivative works. That’s legally complex. At a minimum, understand the foundation model’s licensing terms and what restrictions they place on your use.

The organizations that are handling this well treat training data provenance as seriously as they treat model accuracy. They have someone who owns this question. They document their decisions. They don’t assume they have rights they haven’t verified. They distinguish between “we can use this data” and “we can deploy a model trained on this data.”

The ones getting in trouble treat it as a problem to solve after the model is built, or as something that legal will “figure out” later. They inherit models trained on questionable data. They deploy models trained on that data without understanding the licensing constraints. Then they’re surprised when legal gets involved.

Your model might be excellent. Your training data might be the right data for the problem. But if you can’t defend your rights to use that data, you can’t deploy the model. That’s not a legal problem you solve later. That’s a governance problem you solve now.

In Practice: AI in the Enterprise | Day 23: When AI Breaks Your Workflow (Not Because It’s Bad, But Because You Never Designed for Failure)

I’m sitting across from an operations leader. Their AI system is working—performing exactly as designed. It’s also wrecking their business.

The system makes a recommendation every hour for a critical workflow decision. It’s right about 92% of the time, which is significantly better than humans. But the workflow was built assuming humans would always be there to catch the 8% when it’s wrong. Now the 8% is happening, and there’s no human in the loop. There’s no fallback. There’s no way to pause and escalate. The system keeps recommending and the business keeps following it until someone notices the outcomes are getting worse.

This is not a model problem. It’s a design problem.

I’ve watched this play out differently in three organizations. One built an AI system to route support tickets to the right team. It worked great until the work volume spiked—suddenly tickets were being misrouted to teams that were already underwater, creating cascading delays. The model didn’t fail; the system failed because there was no surge protection. No way to throttle input when the system couldn’t keep up. No circuit breaker. It just kept processing and making worse problems worse.

Another organization integrated AI into their hiring pipeline to screen resumes. The model was well-built, validated, tested. Then they realized halfway through a hiring season that they’d never thought about what happens when the model is confidently wrong about candidates in a particular geographic region. The system had been rejecting qualified people systematically. When the model failed, there was no recovery mechanism because recovery wasn’t part of the design.

A third built an AI system to predict customer churn and trigger retention actions automatically. It worked for six months. Then the retention actions themselves changed customer behavior in ways the model didn’t anticipate. The model started making predictions based on patterns that no longer applied. The system cascaded—triggering increasingly aggressive retention offers because churn was getting worse, which made the problem worse.

These are all preventable disasters. Not because the AI was bad, but because the system design didn’t account for the ways AI could fail.

This is about operational resilience, which is different from accuracy. Accuracy is about the model being right. Resilience is about the system continuing to work when the model is wrong.

Start here: Every AI system needs a failure mode inventory. Not “what if the model is inaccurate?” Everyone assumes that. I mean: What if the model is confidently wrong about a specific type of situation? What if the distribution shifts? What if the system’s own outputs create feedback loops that degrade performance over time? What if volume spikes and latency gets worse? What if the system stops receiving input? For each of these modes, what happens to your business?

Then you design for each one. That means fallback paths. If the model is wrong about X category, what’s the human-validated process that catches it? If volume spikes, what throttles the system to maintain quality? If distribution shifts, what’s the early warning system that alerts you before business impact is severe?

Next: Build in observability from day one. Not for data scientists. For operators. The people running this system in production need to understand, in real time, whether the system is healthy. That means logging which decisions the system made, what inputs drove those decisions, and which ones turned out to be right. It means tracking performance over time—not just overall, but stratified by decision type and context. It means dashboards that show “this part of the system is degrading” before “the whole system is broken.”

Then: Define clear escalation paths. When the system detects it’s entering unknown territory, what happens? Does it keep going? Does it hand off to a human? Does it pause? You need to decide this before it happens. And you need to build it into the system’s logic. A lot of organizations treat escalation as a manual thing. Someone has to notice the anomaly and decide to escalate. By then you’re already in crisis. Better to build escalation rules into the system itself. If confidence drops below X, escalate. If this type of decision appears and we haven’t seen it before, escalate.

Finally: Test failure modes before production. Not happy-path testing. Not “what if data is malformed?” Test the specific failure modes you identified. What happens if the model is consistently wrong about segment B? Can your system detect it? Does it cascade? Can you recover? Run this test. See what breaks. Then fix it.

The reason these disasters happen is that organizations think about AI systems as if they only have one failure mode: the model is wrong. So they focus on model accuracy. But in a production system, there are dozens of failure modes. The model could be right but the training distribution has shifted. The model could be confidently overconfident. The system could be amplifying its own errors through feedback loops. The upstream data could have changed. The business context the model was trained for could have evolved.

These aren’t model problems. They’re system design problems. And they require thinking about the whole flow: inputs, processing, decision, action, feedback.

The organizations doing this well don’t separate AI from the business process. They design the AI as a component of a larger resilient system. That system has fallbacks. It has observability. It has circuit breakers. It has clear escalation rules. It gets tested not just for accuracy but for how it fails.

That’s the difference between building an AI system and building an AI system that doesn’t wreck your business when it inevitably makes mistakes.

In Practice: AI in the Enterprise | Day 22: The Dashboard Your Board Should Be Looking At (Hint: It’s Not What You’re Showing Them)

Your current AI metrics dashboard is probably fine. For a data science team.

It’s terrible for a board.

I sit in governance reviews where the AI lead pulls up a screen showing model accuracy trending upward, latency trending downward, inference volume climbing. It’s all objectively good. The model is getting better, faster, and busier. Then someone asks: “So what could go wrong?” And the room goes quiet because no one actually knows.

This is the measurement problem that organizations skip. They measure what the model does. They rarely measure what the model breaks.

There’s a difference between operational metrics and risk metrics, and most boards are staring at the former while missing the latter. You can have a model that’s technically performing well and operationally failing your business in ways your dashboard doesn’t surface.

Let me break down what I mean. Your current dashboard probably shows something like: Model accuracy (95.2%, up from 94.8% last quarter). False positive rate (12%, down from 14%). Inference time (42ms, stable). Monthly predictions (1.2M, up 18%). These are real metrics. They matter. But they answer a specific question: Is the model working as built?

The board needs to answer a different question: Is the model working for us?

That requires a different set of metrics. Start here: What’s the cost of being wrong? If your model handles credit decisions, false positives cost you acquisition (you reject good customers). False negatives cost you risk (you approve bad ones). Different costs entirely. Your current accuracy metric doesn’t distinguish between them. A board needs to see the business impact ratio. Your false positive rate matters less than your false positive cost. Same with false negatives.

Next: What’s changed about the world your model operates in? This is the distribution shift metric that rarely makes it onto governance dashboards, but it should be top-line. You can track this through several lenses. Are the input features drifting? Are the outcomes changing? Is the relationship between input and outcome degrading? A model that was trained on pre-pandemic customer behavior is still technically operating correctly in 2026—it’s just operating on outdated assumptions. Your board needs to know this is happening before you optimize for patterns that don’t exist anymore.

Third: Where is the model breaking? Accuracy across the full dataset is meaningless if the model is 97% accurate for one customer segment and 72% for another. A board needs to see performance stratified by the segments that matter to your business. This often reveals something uncomfortable: your model is optimized for your largest customer segment and failing your most vulnerable population. That’s a governance problem that a single accuracy number hides completely.

Fourth: What would we know if something went wrong? This sounds abstract, but it’s crucial. If the model’s performance degrades, how quickly do you detect it? What’s your lead time between degradation and detection? If it’s three weeks, you have a problem. A board needs to see mean time to detect (MTTD) as a governance metric. It’s the difference between catching a problem while it’s affecting thousands of customers versus catching it while it’s affecting millions.

Fifth: What’s the audit trail? If a customer or regulator asks why they received a particular decision, can you explain it? Not in terms of model weights or feature importance. In terms of: Here’s what the model saw, here’s how it weighted the information, here’s what it decided, and here’s whether we think that was right. This requires logging and auditability standards that most systems don’t have. A board needs to see what percentage of decisions are explainable within your governance standard. If it’s 85%, you have meaningful risk exposure.

Sixth: How much are we trusting the model? This is the most valuable and most overlooked metric. What percentage of decisions are the model making alone versus in partnership with humans? What percentage of model-only decisions are being overridden by downstream processes? If the model is making recommendations but humans are overriding them 40% of the time, you either have a trust problem or a model problem. Either way, you’re not getting what you paid for. A board needs to see this metric because it shows the difference between model adoption and actual adoption.

Here’s what makes this hard: These metrics require you to actually track things that organizations often treat as incidental. Logging which decisions were right and which were wrong. Tracking demographic breakdowns of model performance. Measuring the time between detection and decision. Documenting override rates. These aren’t sexy dashboards. They’re operational hygiene. But they’re what a governance-minded board actually needs to see.

The organizations that have moved from compliance reporting to genuine risk management have all done the same thing: They stopped thinking about AI metrics as a data science problem and started thinking about them as a business accountability problem. That reframe changes what gets measured.

Your accuracy metric tells you about your model. The metrics I’m describing tell you about your risk. Your board needs both, but if you can only show one, they need the second one. Most organizations show the first and hope no one asks about the second. That’s backwards. Start there.

In Practice: AI in the Enterprise | Day 21: You Can’t Have AI Governance Without Data Governance

The easiest way to spot a governance framework that won’t work is to ask a simple question: “What happens when your training data is wrong?” Watch how long it takes for someone to answer.

Most organizations have governance structures for AI models that exist entirely downstream. They talk about model cards, testing protocols, access controls, deployment gates. These are real things. But they’re built on an assumption that the data feeding the system is adequate. It rarely is.

I’ve watched this play out in three different ways. First, there’s the organization that has excellent data quality standards—but only for their transaction systems. When AI teams pull data for training, they treat it like a novelty. There’s no service level agreement. No validation pipeline. No understanding of what “good” looks like. The AI governance meeting happens monthly. The data quality review never happens at all.

Second, there’s the organization that built data governance years ago around regulatory compliance. Their data lineage tools work great for financial reporting. But AI training pulls from sources those tools never contemplated—third-party datasets, unstructured feedback, behavioral logs. The data steward doesn’t know what’s in production. The AI team doesn’t know where it came from.

Third, and this one is common: the organization that has separate governance structures entirely. Data governance is a technical problem, owned by engineering. AI governance is a risk problem, owned by compliance. They email each other occasionally. The data person doesn’t understand why the model person needs certain attributes. The model person doesn’t understand why the data person won’t merge this dataset from another system. They optimize locally for their respective problems.

Here’s what actually happens in each scenario: Your AI system goes into production. Three months in, a downstream business process starts failing in ways that don’t quite make sense. You dig. You find that the training data included records from a system migration that was supposed to be flagged as an outlier period. Or the historical data contains behavioral patterns that don’t apply anymore. Or someone changed how a critical field was calculated, and that change wasn’t documented, so downstream models are using stale definitions. Or you discover that 15% of your training set was duplicated because someone ran a batch job twice.

By the time you realize this, the model is in production. You’ve built decisions on top of it. Your governance framework has no way to retroactively assess what changed. It has no way to answer: “Should we have caught this?” Because it has no real view of what the data actually was.

Proper data governance for AI means several concrete things. First, your data contracts have to extend all the way upstream to the source. You need to know not just what fields exist, but what they mean, how they’re calculated, when they change. This is harder than it sounds. A “date” field sounds simple until you discover it’s been calculated three different ways in three different legacy systems that were supposed to be consolidated in 2019 but actually weren’t.

Second, you need a testing mindset around data that matches your testing mindset around code. Before a dataset feeds a production model, it needs to pass checks. Are there unexpected nulls? Has the distribution changed? Are there new values that weren’t in training? This should be automated, not a checklist someone marks off once.

Third, you need traceability. Not just “we trained this model on this dataset,” but “this field came from this system, it was transformed by this pipeline, and these validations happened at each step.” When something breaks, you need to know the full lineage fast.

Fourth, you need an owner. Not a committee. Not a shared responsibility. Someone who wakes up in the morning thinking about whether the data feeding your AI systems is trustworthy. That person needs to have standing in governance conversations. When they say a data source isn’t ready, that’s a model-deployment blocker.

The hard part is that data governance doesn’t feel urgent until it is urgent. A model audit finds a problem, and suddenly you’re in crisis mode trying to figure out where your data actually came from. A regulator asks about training set composition, and you realize you don’t have a complete picture. A customer reports they were treated unfairly, and you need to reconstruct what inputs your model saw.

By then, your governance theater is exposed as purely ornamental.

The organizations doing this well don’t separate AI governance from data governance. They treat them as the same problem. The criteria for putting data into production are the same criteria for putting models into production. Same rigor. Same documentation. Same review. Same ownership. When you ask “How do we know this data is good enough?” you get the same answer you’d get if you asked “How do we know this model is good enough?”

That alignment is harder to build than any framework. But it’s the difference between governance that actually catches problems and governance that exists to check a box.

In Practice: AI in the Enterprise | Day 20: Why Centralizing AI Decisions Is a Trap (But Decentralizing Them Is Too)

I visited two companies in the same week. The first had a centralized AI team. All AI decisions went through them. Nothing shipped without approval. The second had distributed AI. Every business unit built its own AI. The centralized company was slow. The distributed company was chaotic.

The problem isn’t centralization or decentralization. It’s that they’re both optimizing for the wrong thing.

The centralized company thought they were optimizing for control. What they got was bottleneck.

The distributed company thought they were optimizing for speed. What they got was risk.

Neither company understood the actual constraint: you need both speed and consistency. And the structure that gives you both of those isn’t about where you put the team. It’s about where you put the decision.

Why Centralization Doesn’t Work

A centralized AI team sounds good in theory. You have one group of experts. They set standards. They approve decisions. Nothing gets shipped that violates the standards.

What actually happens:

The centralized team becomes the bottleneck. Every business unit has to wait for the AI team to say yes. The business unit thinks: “We have a problem we could solve with AI. But it’s going to take six months to get approval.” So they either don’t solve the problem, or they find a workaround that bypasses the central team.

The central team thinks: “We’re protecting the company from bad AI decisions.” What they’re actually doing is training the organization to either wait or work around them.

Also, the centralized team can’t understand every business context. They make rules that work for some problems but not others. They say “we only use approved models” but different business units have radically different requirements. They say “we monitor for fairness” but fairness metrics are different for hiring vs. lending vs. content recommendation.

The company doesn’t end up with strong governance. It ends up with slow governance that’s frequently bypassed.

Why Decentralization Doesn’t Work

A decentralized structure sounds good on paper. Every business unit moves fast. They understand their own problems. They can experiment and iterate.

What actually happens:

Every business unit builds their own version of the same solution. The hiring team builds a resume screener. The recruiting team builds another one. The L&D team builds a third. None of them talk to each other. They’re all making different choices about data, models, monitoring, governance.

One team uses a model that’s biased against a protected class and doesn’t notice because they never thought to check. Another team’s system breaks down in production and they lose customer data because they didn’t think about operations. A third team’s decisions are unexplainable to regulators because they built a black box without documentation.

The company doesn’t end up with speed and flexibility. It ends up with fragmentation and invisible risk.

And the real problem: when something goes wrong in one business unit, the company’s regulatory exposure is enterprise-wide. The hiring team’s bias problem is a company problem. The recruiting team’s data loss is a company problem. The L&D team’s black box is a company problem.

So decentralization feels fast until it’s not. Then it’s expensive.

What Actually Matters Is the Decision Point

The real question isn’t “where does the team sit?” It’s “where does the authority sit?”

And here’s the shift that changes everything: authority should be at the decision point, not at the team location.

Let me explain what I mean.

In a centralized company, authority for all AI decisions is with the central team. A business unit comes and says: “We want to build a hiring screener.” The central team decides. That’s the problem. The central team doesn’t know enough about hiring to make that decision well.

In a decentralized company, authority is with the business unit. They decide. But the company has no way to ensure consistency or manage enterprise risk. A business unit makes a decision that creates regulatory exposure for the whole company, and the company can’t prevent it.

What actually works is splitting the authority:

The business unit decides: “We have a problem and we think AI can solve it.” That’s their decision. They know their problem better than anyone. They should have authority here.

A shared standards team flags: “If you use AI this way, here’s what you need to do around data quality, fairness, explainability, operational risk.” These aren’t approvals. They’re constraints.

The business unit decides: “We’ll meet these constraints with approach X.” That’s their decision. They know how to do it.

A shared audit function checks: “Are you actually meeting the constraints you committed to?” They’re not blocking. They’re verifying.

This distribution of authority gives you what you actually want:

  • Speed: business units can move fast because they’re not waiting for approval
  • Consistency: everyone is building to the same enterprise constraints
  • Risk management: you’re checking for problems without blocking problems
  • Accountability: it’s clear who decided what, and it’s clear why it was decided

How Standards Without Approval Actually Works

The shift from “approval authority” to “standard-setting authority” is subtle but important.

Approval means: “We get to decide if you can do this.” You’re slowing down the organization.

Standards means: “Here are the constraints. Meet them, and you’re free to move.” You’re enabling the organization.

The centralized company is doing approval. The distributed company is ignoring standards. Neither is working.

What works is: clear standards that business units must meet, clear authority for business units to decide how to meet them, and clear verification that they’re actually meeting them.

Here’s what this looks like in practice:

A business unit wants to build a pricing model. The standards say: “If you use AI for pricing decisions, you need to: – Document the data sources and any known biases – Verify that recommendations don’t systematically harm specific customer segments – Have a manual override process for edge cases – Monitor recommendation outcomes against actual pricing decisions – Review outcomes quarterly with your business leader and the standards team”

The business unit says: “We’ll do all that.” Great. They don’t need approval. They can start building.

Six months later, they’re in production. The standards team reviews: “Are they actually documenting data? Are they actually monitoring outcomes? Are they actually doing quarterly reviews?” Yes? Great. No? Then you have a conversation about what’s getting in the way.

The business unit didn’t get held up by approval. They also didn’t go rogue. They built to standards.

The Scaling Problem

The reason this matters is that your company is going to deploy AI across every business unit. Not maybe. It’s happening.

If you centralize, you’ll be the slowest company in your industry because you’re approving every experiment, every iteration, every new use case. Your competitors will be 10x ahead of you.

If you decentralize without standards, you’ll have governance risk in every unit, security risk in most of them, and regulatory exposure enterprise-wide.

The middle path is: clear standards, distributed authority, shared verification.

This requires:

  1. Someone owns the standards. Not approves. Owns. Writes them. Updates them based on experience. Makes sure they’re actually good standards, not theater standards.

  2. Business units have clear authority. They decide to use AI. They decide how to meet the standards. They move fast. They’re not waiting for approval.

  3. Verification is real. Someone’s actually checking if the standards are being met. Not by reviewing documents. By checking outcomes, talking to teams, monitoring in production.

  4. Escalation is clear. If a business unit can’t meet a standard, there’s a clear escalation. Not “you’re blocked.” But “here’s the constraint, let’s figure out how to address it.”

The Mistake That Kills This Approach

Most companies can’t do this because they try to start with standards that are too prescriptive.

They say: “All AI systems must use approved models.”

That’s not a standard. That’s an approval requirement hidden in standard language. And business units will either ignore it or spend six months trying to get a model approved.

Real standards look like:

  • “All AI systems must have a documented rationale for the data sources.”
  • “All AI systems used for consequential decisions must have outcome monitoring.”
  • “All AI systems must have a person accountable for outcomes.”
  • “All AI systems must have an escalation process if they start making decisions outside their training distribution.”

These are standards that you can verify. They’re not gatekeeping. They’re clarifying what good looks like.

The Organizational Structure That Enables This

You don’t necessarily need to reorganize. But you need clarity about authority:

Business unit leaders decide what problems to solve with AI. They have authority. They move fast.

A shared standards function (could be central AI team, could be part of Risk, could be part of Architecture) owns standards, not approvals. They set expectations. They verify. They escalate when needed.

An audit or verification function checks that standards are actually being met. This could be internal audit, compliance, or a dedicated team.

Escalation authority (could be CTO, CRO, business leader, doesn’t matter who) is clear. When a standard conflicts with a business need, there’s a clear path to resolve it.

This looks different from a pure centralization or pure decentralization. It’s designed for the actual constraint: you need speed and consistency simultaneously.

What You Should Do Monday Morning

If your company has a centralization-decentralization argument going, the question to ask is: “What are we actually optimizing for?”

If the answer is “we want to move fast,” decentralization feels right until you hit risk problems.

If the answer is “we want to manage risk,” centralization feels right until you slow down the business.

If the answer is “we want to move fast AND manage risk,” then the answer isn’t about where to put the team. It’s about where to put the authority.

Ask your leadership: “Where is authority for AI decisions in our company? Is it business units? Is it the central team? Is it clear?”

If you can’t answer in a sentence, you have a structural problem.

The companies that are scaling AI aren’t solving this with better committees or more governance layers. They’re solving it by being clear about where decisions are made and why.

Centralize authority where you need consistency. Decentralize authority where you need speed. But be explicit about where each one lives.

In Practice: AI in the Enterprise | Day 19: What Enterprises Discover When Regulators Examine Their AI Systems

A regulator walked into a bank’s AI governance office and asked a simple question: “Show me your decision log for the AI systems you’re using in lending.”

The bank produced a spreadsheet: model name, version, accuracy metrics, deployment date.

The regulator asked: “Where’s the decision log? When did you decide this model’s output? When did it change? Why did you decide to use this model for this decision? Who approved it? What was the alternative?”

Blank looks.

What the bank had was a model registry. What the regulator was looking for was a decision record.

And the reason this matters is that the regulator’s job isn’t to check if your model is accurate. It’s to check if you made a reasonable decision about how to use AI in your business, and you can defend that decision.

Most enterprises aren’t prepared for that conversation. Not because they’re incompetent. But because they’ve been building technical infrastructure (governance committees, model monitoring, fairness frameworks) when they should have been building decision infrastructure.

Why Regulators Care About Decision, Not Accuracy

Here’s the shift that’s happening in regulatory thinking, and most enterprises haven’t noticed it yet:

Regulators care about what you decided to do with AI, not whether you did it well.

If you can show: “We decided to use AI in this part of our business because X. We evaluated the alternatives. We identified the risks. We have governance to manage those risks. We monitor outcomes. We’ve decided to accept these specific risks. Here’s how we track if we’re right.”

Then the regulator might actually approve. They might ask hard questions, but you have a story.

If you can’t show that—if instead you show: “We built a really good model. Look at the accuracy. Look at the fairness metrics.”—you’re not answering the question the regulator is asking.

The regulator already assumes your model is good. What they want to know is: Why did you use AI here? What problems could this cause? Have you thought about it? Who’s accountable if something goes wrong?

Most enterprises have built answer to “Is the model good?” They haven’t built an answer to “Did you decide to use AI responsibly?”

These are different questions, and they require different infrastructure.

What the Regulator is Actually Checking

When a regulator looks at your AI system, they’re asking:

Decision authority: Who decided to use AI here? Was that person qualified to make this decision? Did they understand the risks?

Alternatives considered: Did you evaluate other ways to solve this problem? Why did you choose AI? Why not just use rules? Why not keep the human in the loop?

Risk assessment: Did you identify what could go wrong? Have you thought about fairness, accuracy, security, operational risk? Or did you just think “this model is accurate”?

Monitoring and accountability: How will you know if something goes wrong? Who’s accountable? Do you have a way to fix it if you need to?

Documented reasoning: Can you show me the decision? Not the model performance. The decision about whether to use the model.

Most enterprises would fail this audit today. Not because your models aren’t good. But because you can’t show the decision.

You have model cards. You don’t have decision records.

What You Actually Need to Build

Here’s what “decision infrastructure” looks like:

Decision record template. Before you deploy an AI system, someone fills this out: What decision are we using AI to make? What are the alternatives? Why are we choosing AI? What could go wrong? What’s our monitoring plan? Who’s accountable?

This is boring and unsexy. Nobody gets excited about decision records. But when a regulator asks “how did you decide to use AI here?” you can point to a document that answers the question.

Audit trail. When you change a model, when you change monitoring, when you escalate a risk, you should have a record. Not “we deployed version 2 on this date.” But “we decided to change the monitoring thresholds because X, here’s who approved it, here’s the reasoning.”

This isn’t new. Banks have done this for lending decisions for decades. The infrastructure exists. Most enterprises just haven’t applied it to AI systems.

Escalation record. When something goes wrong with the AI system (accuracy drops, outcomes drift, fairness metric breaks), do you have a record of what happened and what you decided to do? That’s what regulators care about. Not that nothing went wrong. But that you caught it and did something about it.

Most enterprises have an incident response plan. They don’t have a “what did we decide when we found a problem with the AI system” plan.

Approval authority. When you want to deploy an AI system or make a material change to one, who approves it? Not a committee. A person. With clear authority. Who will own the decision if something goes wrong.

The Conversation Regulators Are Having

Regulators are starting to enforce this. It’s not consistent yet (different regulators care about different things), but the pattern is becoming clear:

“You have to be able to explain your AI decisions.”

This is not “you have to have a fair model.” This is “you have to be able to explain why you decided to use a model, what you evaluated, and what you’re doing to make sure it works.”

Some regulators care more about fairness. Some care more about security. Some care about operational risk. But all of them care about the decision.

And they’re discovering what I’ve seen in every enterprise: most companies can’t explain the decision. They can explain the model. The model is great. But the decision to use the model in the business? That decision was never explicitly made.

What This Means for Your Enterprise Now

If you’re waiting for regulatory guidance to clarify what they want, you’re late. Regulators are starting to conduct reviews now, and they’re finding the infrastructure gap.

What you should do:

  1. Start with one model. Pick a model that’s already deployed. Walk through the decision. Who decided to use it? Did they document their reasoning? Do you have an alternatives analysis?

  2. Build the decision record backward. You don’t have to have perfect decisions. But you have to be able to explain them. When a regulator asks “why did you use AI here?” you should be able to answer in a sentence or two.

  3. Set up escalation tracking. When something changes with the model—accuracy drops, fairness metrics break, operational impact—do you have a record of what you discovered and what you decided to do?

  4. Identify your approval authority. This is a person, not a committee. They can consult the committee. But they own the decision.

  5. Plan for the hard conversations. When a regulator asks “can you defend this decision?” you might realize you can’t. That’s okay. You’re learning. But it’s better to learn this before a formal review.

The Difference Between Compliance Theater and Actual Compliance

Here’s the trap: you can build all the governance theater in the world—committees, frameworks, monitoring—and still not be able to answer a regulator’s actual questions.

Real compliance means: when a regulator asks “did you think about this?” you can say “yes, here’s what we thought, here’s what we decided, here’s how we’re managing it.”

Theater means: when a regulator asks “did you think about this?” you produce a document that says “we have governance” but doesn’t show that you actually thought about anything.

The companies that are handling this well are building decision infrastructure, not governance infrastructure. They’re keeping records of decisions. They’re being explicit about who decided what. They’re documenting their reasoning.

It’s not impressive. It’s not technically sophisticated. It’s boring process work.

But when a regulator asks for it, you have it. And that’s the difference between a review that goes smoothly and one that doesn’t.

The Regulatory Reckoning

There’s a regulatory reckoning coming. Not because regulators are anti-AI. But because regulators are starting to apply the same bar to AI systems that they apply to everything else: Can you defend the decision?

Most enterprises would fail that bar right now.

The time to fix this is now, not when you’re in a regulatory review. Build the decision infrastructure. Document your thinking. Make explicit decisions with clear authority.

Because when a regulator asks “show me your decision log,” you’re going to want to have one.

In Practice: AI in the Enterprise | Day 18: Model Risk Looks Different in Enterprises (And You’re Probably Using Consumer-Grade Frameworks)

A bank’s Chief Risk Officer pulled out a model risk framework. It was comprehensive: data quality checks, concept drift monitoring, backtesting protocols, confidence intervals, stress testing. It looked good. Very academic.

Then I asked: “What happens when a model decides to drop a product line that’s currently profitable but the model thinks is risky?”

Silence.

“What if the model’s risk assessment is more conservative than the business risk tolerance, and nobody knows that’s the conflict?”

More silence.

“Who resolves that conflict?”

This is the gap that almost every enterprise is walking into: they’re using risk frameworks designed for research or trading (where “maximize accuracy” is the goal) and applying them to production systems in regulated industries (where “make the right business decision” is the goal, and accuracy is just one input).

These are not the same problem.

Why Academic Frameworks Don’t Translate

Let me be clear about what academic model risk frameworks do well:

They identify technical failure modes. They catch when training data has changed. They detect when model assumptions are violated. They flag when a model’s confidence is misplaced. These are real, important things.

What they don’t do is address business risk. They can’t, because business risk is not the same as model risk.

In a research setting, a model is either right or wrong. In an enterprise setting, a model is making a decision that has business consequences. The risk isn’t “the model is inaccurate.” The risk is “the model is making a decision that conflicts with our business goals or regulatory constraints, and we didn’t catch it.”

Here’s the concrete difference:

A research model that fails: “The accuracy dropped from 94% to 91%. Something is wrong with the model.”

An enterprise model that fails: “The model is rejecting 15% more credit applications than it used to. Is that because credit quality changed, or because we changed the training data, or because the model is malfunctioning, or because the model’s risk tolerance diverged from the business risk tolerance?”

The research framework will catch some of this (data changes, concept drift). It won’t catch the one that matters most: we didn’t check if the model’s decision-making is aligned with what the business actually wants.

The Misalignment Nobody Talks About

Here’s what I see in most enterprise risk frameworks:

“Monitor model performance. Flag degradation. Retrain. Validate. Deploy.”

What’s missing: “Validate that the model’s decisions are aligned with business goals.”

And the reason it’s missing is that people don’t think of it as a model problem. They think of it as solved by governance. “We have a governance committee that approves what the model does.”

But the committee approves the model. They don’t continuously validate that the model is doing what it approved.

So you get scenarios like this:

A model for employee promotion recommendations gets deployed. The governance committee approved it. The accuracy is good. The fairness testing passed. It’s in production.

Six months later, someone notices the model is recommending promotions for technical specialists into management roles at much higher rates than the business actually wants. The model learned that technical expertise is predictive of success, which is true statistically. But the business doesn’t want to promote all the good engineers into management. That’s not what the business risk tolerance is.

Is this a model risk problem? In the academic sense, no. The model is doing exactly what it was trained to do. In the business sense, yes. The model’s decisions conflict with what the business actually wants.

The academic framework checks: “Is the model accurate?” Yes. “Is the model fair?” Yes. “Is the model stable?” Yes. “Is the model doing what the business wants?” Never asked.

The Risk Framework You Actually Need

Real enterprise model risk has three layers:

Layer 1: Technical Risk. Is the model working as designed? This is what academic frameworks address. Data quality, concept drift, distribution shift, confidence calibration. These matter. Monitoring these is necessary but not sufficient.

Layer 2: Business Risk. Is the model’s decision-making aligned with business goals? This is where most frameworks stop. You need to actually validate that the model is making the business decisions you want it to make, not just making accurate predictions.

Layer 3: Governance Risk. Is someone explicitly accountable for what the model does? And if the model’s behavior conflicts with policy or regulation, do you have a way to know and a way to respond?

Here’s how you’d actually monitor a credit model:

Technical layer: Monitor data quality (are the features that go into the model staying the same?). Monitor performance (is accuracy stable?). Monitor distribution (are the inputs to the model changing in unexpected ways?).

Business layer: Monitor outcomes in context. If approval rates are changing, is that because credit quality changed, or because the model changed? If the model is approving more customers from some segments than others, is that intentional? What’s the approval rate spread you’re willing to accept?

Governance layer: Is there someone who’s reviewed the outcomes and said “yes, this is what we want”? If the model is declining business that was profitable, who made that decision? If the model’s risk tolerance is more conservative than the business wants, who’s changing which one?

Why This Actually Matters for Your Business

The reason this distinction is critical is that most enterprise models are making proxy decisions.

You don’t actually want a credit model that’s “accurate at predicting default.” You want a credit model that makes profitable lending decisions, recognizes systemic risk, and complies with regulation. Accuracy is a means to those ends, not the goal itself.

If your model is 95% accurate at predicting default but it’s also declining 30% of applications that would have been profitable, you have a problem. The academic framework will say “the model is accurate.” The business framework will say “the model is costing us money.”

In a lot of enterprises right now, this is exactly what’s happening. Models that were built to be accurate are getting used to make business decisions, and nobody’s checking whether the business decisions are good.

The technical team builds a model that predicts well on the training data. It passes validation. It gets deployed. The business starts using it. Six months later, someone notices it’s making different decisions than the human process it replaced, and nobody knows if that’s good or bad.

What Changes When You Separate These Risks

The first thing that changes is accountability. If a model is making bad business decisions, it’s not “a model accuracy problem.” It’s “a governance problem” or “a business decision problem.” And the person who’s accountable is not the data scientist. It’s the person who’s using the model to make decisions.

The second thing that changes is what you monitor. You don’t just look at technical metrics. You look at decision outcomes. You build feedback loops that tell you whether the model is doing what the business wants it to do.

The third thing that changes is escalation. If the model is making technically correct predictions but bad business decisions, you need a clear path to change the model, change the policy, or change the usage. And you need someone accountable for that decision.

The Conversation to Have Now

Most enterprises have started monitoring technical model risk. That’s good. But they haven’t started monitoring business model risk, and they definitely haven’t integrated that into governance.

The conversation you should be having is not: “Do we have good model monitoring?” It’s: “Do we know whether our models are making the business decisions we want them to make?”

That requires:

  1. Clear business objectives for each model. Not “predict well.” What decision do we want the model to make? What constraints does it have? What risks are we accepting?

  2. Outcome monitoring that’s connected to business metrics. If the model changes its behavior, do we know? If the decision outcomes change, is that intentional?

  3. Escalation clarity. If the model’s decisions diverge from the business goals, who has authority to change the model or change the policy?

  4. Explicit ownership. Not “the data science team owns the model.” Someone owns “we’re using this model to make X decision and it’s working as intended.”

This is not about building more sophisticated risk frameworks. It’s about connecting the technical risk work to the business decision that’s actually being made.

The Risk That Looks Like Success

The most dangerous scenario is the one where the model is technically sound, it’s making consistent decisions, it passes all the monitoring checks—and it’s making decisions the business doesn’t actually want.

You won’t notice this until it matters. Until someone asks, “Why did we decline that credit application?” and the answer is “the model said so,” and the question back is “but is that what we actually want?”

If you can’t answer that quickly and clearly, you have a model risk problem that no amount of academic monitoring will catch.

In Practice: AI in the Enterprise | Day 17: The Accountability Inversion: Why Blaming the Data Team Is Convenient and Wrong

A financial services company shipped a credit decisioning model that started rejecting credit-worthy applicants from a specific zip code. When the incident got escalated, the conversation went like this:

“How did this happen?”

“The data was biased.”

“Why was the data biased?”

“The historical training data reflected past lending patterns.”

“Who owns the data?”

“The data team.”

So the data team got blamed. The data scientist who built the model got blamed. The incident was closed as a “data quality issue.”

But that’s not what happened. What happened was a business decision: to use historical lending data to train a credit model, knowing that the data contained biases from past discrimination. And nobody made that decision consciously. Nobody even knew it was a decision.

This is the accountability inversion. We blame the people who execute (the data team) and ignore the people who decided what to build (everybody else). And we do this systematically, in almost every company I’ve worked with.

How the Inversion Works

Here’s the mechanism. You have a process: data sourcing → data cleaning → model training → validation → deployment.

Someone (product, business, leadership) decides what problem to solve. Someone in product/engineering decides how to approach it. The data team gets requirements. They execute against those requirements. A model comes out. It gets shipped.

Something goes wrong.

Everyone looks back at the process. The data team was the last team that touched the model before it broke. So the data team gets blamed.

But the failure wasn’t in execution. The failure was in the decision about what to build.

If the requirements were: “Build a model using historical lending data to predict credit-worthiness,” and the data team said, “This data contains systemic bias,” and nobody responded by changing the requirements, then the person who ignored that flag made a decision. They decided to accept the risk of bias to get a faster model. That’s a legitimate decision. But it’s not the data team’s decision. It’s the business’s decision.

Instead, what happens is the data team raises the concern, the conversation stops, and when the model breaks, people say: “The data was biased.” As if bias is a natural property of data rather than a decision about which data to use.

The accountability inversion happens because people at the decision point don’t see it as a decision. They see it as “running the process.” The data team sees it as a constraint. And when something goes wrong, the constraint gets blamed.

Why This Matters More Than You Think

This matters because companies optimize based on where they assign blame.

If you blame the data team for bias, your company will hire better data scientists and invest in better data tools and build better data governance. These are all good things, but they don’t prevent the problem you actually had.

The problem you had was: someone built a model that discriminates against a protected class. That’s not a data problem. That’s a decision problem. That’s a product decision, a business decision, and a risk decision.

If you blame the data team, you’re saying: “The process failed because the execution wasn’t good enough.” So you make the execution better.

If you actually solved the problem, you’d be saying: “The decision-making was wrong. We made a conscious choice to use biased data and didn’t have the governance to catch it before shipping.” So you change the decision-making, not the execution.

And this inverts the incentives. The data team, knowing they’ll be blamed if something goes wrong, becomes more conservative. They push back harder on any requirements that seem risky. The business becomes more frustrated with data teams, because they slow everything down. The relationship becomes adversarial. And you end up with governance theater again.

Meanwhile, the actual decision-makers feel less pressure to think carefully about what they’re building. The worst outcome gets caught by the data team, so their job is to avoid shipping anything the data team flags. That’s a low bar.

What Actually Went Wrong

Let me walk through what should have happened in the credit model case.

Product says: “We want to predict credit-worthiness.”

Engineering/Data says: “We can do that. Here are the approaches: – Option A: Use historical lending data (fast, cheap, but contains systemic biases from past discrimination) – Option B: Use transactional data and alternative credit indicators (slower, more expensive, but less biased) – Option C: Use only recent lending data and accept that we have limited historical performance (somewhere in between)”

Then someone (product, business, risk, compliance) decides which option. That’s the decision. That’s the point where someone is making a choice about which risks to accept.

If you choose Option A, the decision is: “We’re accepting the risk of discriminatory outcomes in exchange for speed and cost.”

That’s a legitimate choice. You might make it. But it’s not a hidden choice. It’s an explicit decision that someone needs to own.

What usually happens instead is that the decision gets bundled into “let’s build the model” and nobody explicitly decides which option to take. The data team gets some vague requirements, picks the fastest path, and builds the model. When it breaks, you blame the execution.

The Three People Who Made the Decision

Here’s who actually decided to accept the bias risk:

  1. The product person who framed the problem. If they framed it as “predict credit-worthiness from historical data,” they’ve already chosen Option A. They made a decision about data sources without realizing it.

  2. The engineering/data person who translated the requirement into a concrete approach. They could have pushed back and said “this data is biased; here are the implications.” If they didn’t push back, they made a decision to accept it. If they did push back and were ignored, then the person who ignored them made the decision.

  3. The business/leadership person who accepted the timeline and cost of Option A. If you decide your risk tolerance is “whatever the data allows” rather than “we need to actively manage bias,” that’s a decision with real implications.

Nobody is evil here. Nobody is incompetent. But three different people made pieces of a decision without any of them explicitly owning the whole decision. So when the model breaks, everyone points at someone else.

The data team gets blamed because they’re the last person in the chain. But they didn’t decide to use biased data. They were given a requirement and executed it.

What Accountability Actually Looks Like

Accountability means: someone explicitly decided and someone explicitly owns the outcome.

In the credit model case, here’s how it should work:

Before shipping: – Product frames the problem. (Decision 1: what are we solving?) – Product/Engineering present options with tradeoffs. (Transparency: here’s what each option actually means) – Someone (let’s call it Risk or Product leadership) decides which option. (Decision 2: which risk do we accept?) – Data/Engineering executes to that decision. (Execution with clear constraints)

If something goes wrong: – If the model breaks because of a data quality issue (execution failure), the Data team owns it. – If the model breaks because of a decision we made (the bias was predictable from the approach), the person who made the decision owns it.

The accountability framework clarifies who decided what. It doesn’t blame people for circumstances they didn’t choose. But it also doesn’t let decision-makers hide behind “the data was biased.”

The Rationalization That Kills Accountability

I’ve watched this happen in almost every company:

Something goes wrong. Someone asks: “Whose fault is this?”

The convenient answer is always “the process failed” or “the data was bad” or “the execution wasn’t good enough.”

The hard answer is: “We made a decision about what to build that we should have thought through more carefully.”

So companies invest in better processes, better data tools, better execution frameworks. And the same decision-making failures happen again, just with better process theater around them.

What changes when you flip the accountability is that decision-makers have to think about what they’re actually deciding. It’s not: “Do we want to move fast?” It’s: “If we move fast using this approach, what are the actual implications? Who’s the decision owner if those implications are bad?”

That conversation is uncomfortable. It’s easier to blame the data team.

The Pattern to Watch For

If your company has a pattern where the data team is blamed when ML models break, or where technical teams are blamed for bad outcomes, or where execution is blamed for strategy failures, you have an accountability inversion.

The fix isn’t better data governance or better technical execution (though those might help). The fix is clarity about who decided what.

Stop asking: “What went wrong with the execution?”

Start asking: “What decision was made about what to build, and who made it?”

And when you get that clear, you stop blaming the data team. You either blame the person who made the decision badly, or you stop treating it as a blame situation and start treating it as a learning situation.

Actually, you probably want to stop treating things as blame situations anyway. But first you have to stop blaming the wrong people.