In Practice: The Other Side of the Table | Lesson 2: The Problem You Can Actually Name

You have a mandate, a budget line, and a category. What you do not have, yet, is a sentence describing what breaks if you do nothing.

Almost every AI evaluation I have watched start badly started here. Not with a bad problem. With no problem, and a capability standing in for one.

Capability wearing a problem’s coat

The requirement arrives phrased as an absence. We don’t have an assistant for the service team. We have no way to summarize case notes. We’re behind on document extraction. Each of those describes a thing the company lacks, which feels like a problem and is not one. A company lacks thousands of things and is fine.

The reason this matters is not philosophical. It’s that a vaguely stated requirement cannot discriminate between vendors, and an evaluation that cannot discriminate will be settled by something else. Usually the demo. Sometimes the relationship. Occasionally the price, which sounds responsible and isn’t, because you’ll have picked the cheapest version of a thing you never established you needed.

Take a composite drawn from a few consumer-facing evaluations: a European retailer, contact center, roughly 900 agents. The stated requirement was AI in customer service. Six vendors were invited. The RFP ran to 140 requirements, most of them inherited from the CRM purchase three years earlier and updated by find-and-replace.

What came back were six demos of six different products, each of which met the requirement as written, because the requirement as written could be met by nearly anything. One showed agent-assist suggestions. One showed full deflection at the front door. One showed post-call summarization and quality scoring, which is a different budget, a different owner, and arguably a different project.

The evaluation team spent a month building a comparison matrix before anyone said out loud that the six products were not substitutes for one another. That month was not wasted exactly. It just bought a piece of information they could have had for free in week one.

The sentence

The test I use now takes about twenty minutes and it is uncomfortable in a useful way. Write one sentence in this shape, filling in every slot:

Today, [named group] spends [measured quantity] doing [specific activity], which costs us [a number], and I know that because [named source].

Four of those slots are easy. The fifth is the one that does the work.

If the last clause is “because the team says so,” you have an impression. Impressions are worth investigating and are not worth a procurement. If the last clause names a system, a report, a time-and-motion study, a ticket volume export, then you have a problem, and a problem can be quoted against.

Three ways the sentence usually fails, all of them recoverable.

It names a capability, not an activity. “We lack automated summarization” is not an activity anyone spends time on. Rewrite until the subject of the sentence is a person doing something.

It names a metric with no owner. Reduce average handling time by twenty percent is a fine ambition and a poor requirement, because handling time already belongs to somebody. Go and ask that person whether they think it’s the problem. Sometimes they’ll tell you handling time is up because of a policy change in claims and no software will touch it. That conversation is worth more than the next two months of vendor calls.

It’s three problems in a coat. Deflection, agent assist, and quality scoring share a budget line and nothing else. They have different users, different failure modes, and different vendors. If your sentence needs an “and,” you may have two evaluations.

What breaks if nothing happens

The second test is shorter. Finish this: if we do not do this at all, in eighteen months the consequence is ___.

Answers that survive contact:

  • The backlog reaches a size where we must hire, at an illustrative twelve to fifteen additional heads
  • The regulator’s next thematic review will ask a question we can’t currently answer
  • The contract with our outsourcer renews at a rate we can’t absorb
  • We lose the two people who currently do this by hand, and the process is in their heads

Answers that do not survive contact:

  • We fall behind
  • Competitors will have this
  • The board expects progress on AI

I want to be careful here, because the third one is sometimes the real reason and pretending otherwise helps nobody. A board mandate is a legitimate constraint. It just isn’t a problem statement, and if it’s the actual driver then the honest evaluation is a different one: what is the cheapest credible thing that satisfies the mandate and does no damage. That’s a reasonable question. It leads to a much smaller purchase than the one you’re currently scoping.

Take it to the person whose number moves

Once you have the sentence, it needs one signature, and not from your sponsor.

Find the person who owns the number in the sentence. The service director whose handling time it is. The operations manager whose backlog it is. Show them the sentence and ask two things: is this true, and is this the thing you’d fix first.

You’ll get one of three answers.

They agree, in which case you have an ally who will still be there in month nine when the sponsor has moved on.

They disagree about the number, which means the number is wrong or the measurement is, and you have found this out in week two rather than in the benefits review.

They agree the number is right and tell you it isn’t the thing they’d fix first. This is the most valuable answer and the one people most want to ignore. It usually means the problem you were handed is a symptom, and the tool you’re evaluating will be measured against a cause it doesn’t touch.

I have overridden that answer before, on the grounds that the mandate was the mandate. The deployment worked and the number didn’t move, because the number was being held up by something else entirely. Nobody got fired. It just quietly became the project nobody references.

One thing to do differently

Before you shortlist anyone, get your one sentence agreed in writing by the person whose number appears in it. Not your sponsor, not your steering group. The owner of the metric.

It takes a week. It will occasionally kill your project, which is the point, and it’ll do it while killing the project is still free.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.