In Practice: AI in the Enterprise | Day 8: The Bank That Realized (Too Late) Their AI System Was Making $2M Mistakes in Production

It was a Wednesday afternoon in late October when someone noticed something odd.

A risk officer was reviewing a batch of decisions made by a lending recommendation system. The system had been live for eight months. It was working. Models had been validated. Metrics looked good. But in this particular batch, something didn’t add up.

There was a pattern of approvals that shouldn’t have been approved.

The system had learned to optimize for one thing (approval velocity), and in doing so it had learned to weight certain borrower profiles in a way that made internal sense but created external risk. It wasn’t broken. It wasn’t defrauding anyone. It was just… quietly making decisions that looked right by the metrics the organization was monitoring, but created risk the organization hadn’t accounted for.

When they dug in, they found that over eight months, this system had recommended approval decisions on credits that had characteristics which, under scrutiny, suggested higher default probability than historical norms. The organization had essentially been taking on tail risk that nobody had measured because nobody was measuring the right thing.

The cost, when they finally calculated it, was in the millions. Not because the model was wrong. But because the operational reality of how the system was being used diverged from how it was supposed to be used.

This is not a story about a specific bank. But it’s a pattern I’ve seen map across different domains, different risk types, different sizes of organizations. And it illustrates something fundamental that I think enterprise leaders chronically underestimate: operational risk in AI systems is different from model risk.

Model risk is the question: does the model predict accurately? You can measure that. You can validate it. You can test it on hold-out data.

Operational risk is the question: what actually happens when humans and machines make decisions together in the real world, over time, at scale?

It’s much harder to measure. And it’s where the real money goes.

Here’s how it usually happens:

The system is deployed. It works. Engineers and data scientists are satisfied because the metrics are holding. The business is satisfied because the system is processing decisions faster. Everyone feels good.

Then, slowly, a gap opens up between how the system is being used and how it was designed to be used.

Maybe a human decision-maker starts trusting it too much and stops applying their judgment. Maybe they start using it in cases it wasn’t designed for. Maybe the process changes and the system isn’t updated. Maybe the data it’s learning from drifts because the business has shifted and nobody told the model.

These aren’t failures. They’re just reality. Real humans, real workflows, real organizations. Things change. People adapt. Systems… sometimes don’t.

And in the gap between how the system was designed and how it’s actually being used, risk accumulates. Silently.

The lending scenario is clean because the risk is quantifiable. But the pattern is broader.

A claims adjudication system that learned to approve claims faster starts approving claims that shouldn’t be approved. A hiring system that optimizes for filling positions quickly learns to screen out categories of applicants that were never the intent. A fraud detection system that minimizes false positives starts letting real fraud through. A customer retention system that minimizes churn learns to make offers that are economically senseless.

None of these happen because the model broke. They happen because the operational reality diverged from the model’s intent.

So how do you prevent it?

The honest answer is: you can’t, completely. Humans are adaptive. Systems are rigid. Gaps will open.

But you can reduce it dramatically by doing something most organizations don’t do: operationalize decision monitoring.

Not model monitoring. Decision monitoring.

Model monitoring is: is the model drifting? Is the data quality declining? These are good questions. Most enterprises do some version of this.

Decision monitoring is different: what decisions is the system actually making? Are they consistent with what we expected? When the system and humans disagree, what’s the pattern? What are we learning about how this system is actually being used?

This requires infrastructure that most organizations don’t have. You need to log not just the model’s prediction, but the human’s decision. You need to understand the gap between them. You need to understand how that gap is changing over time.

Then you need someone—an actual human with authority—reviewing this regularly. Not algorithmically. Manually. Reading through cases. Asking: does this feel right? Where should I worry?

The lending organization I referenced earlier, once they figured out what had happened, implemented this. They assigned a risk officer to review 50 random decisions per week from the system. Just reading them. Just asking: do these approvals make sense? Do they align with our risk appetite?

That single investment—one person’s time—would have caught the drift before it became millions in tail risk.

Most organizations won’t do this. It feels expensive. It feels like overhead. The system is working. The metrics are good. Why do you need a human reading cases?

Because operational risk isn’t in the metrics. It’s in the gap between what you’re measuring and what’s actually happening.

Here’s the hard part: you probably can’t see that gap from the metrics. You have to see it by looking at actual decisions, actual outcomes, actual patterns in how the system is behaving in the world.

So if you deploy an AI system that makes real decisions—lending, hiring, fraud, claims, pricing—you need someone whose job is to catch the operational risk. Not to prevent the model from drifting. To catch the moment when the organization’s use of the system starts diverging from the system’s design.

It’s boring work. It won’t show up in quarterly metrics. It’s also how you avoid being the organization that discovers, in month eight, that you’ve been taking risks you didn’t know about.

The cost of operational risk isn’t usually in one bad decision. It’s in the pattern of decisions that are individually defensible but collectively add up to exposure you didn’t want.

Pay attention to that gap. Assign someone to watch it. It’s cheaper than learning the lesson the hard way.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.