In Practice: The Other Side of the Table | Lesson 15: You Are Building a Capability, Not Running a Purchase

The evaluation closes, the contract gets filed, the working group dissolves, and somebody sends a thank you note. Seven months later a different part of the organization starts the same process from nothing.

Nine months, then seven weeks

Composite from a few organizations now several purchases deep: a construction and infrastructure group, on its fourth evaluation of this kind in three years.

Illustratively, the first took about nine months from mandate to signature. The fourth took roughly seven weeks, for a purchase of comparable size and rather more complexity.

The software had not got simpler. What had changed was that the fourth evaluation started with things the first three had produced.

A standing question set, the same one every time. A habit of measuring the baseline before anyone contacted a vendor. A list of the internal groups who could stop a project, with their actual lead times next to their names, updated after each purchase. A rough template for estimating switching cost. And a record of what they had declined to buy, with the reason.

That last one is the rarest. Most organizations keep a careful record of what they bought and no record at all of what they turned down, so the same product gets evaluated twice by different teams and the same objection gets rediscovered from scratch, at full price both times.

What compounds

Four things, and none of them require a center of excellence, a new function, or anybody’s headcount.

Questions rather than templates. A template gets filled in. A question gets answered, and the answers are what you actually needed. The demo requests, the reference questions, the diligence conversation, the pricing question about what moves when usage doubles. Twenty questions on one page, refined a little each time.

The measurement habit. Baseline before vendors. It takes a week and it is the difference between a benefits review that confirms your judgment and one that finds nothing to look at.

The stopper map with lead times. Not a stakeholder list. The groups who can quietly cost you months, and how many months each of them currently costs. This information already exists inside your organization, distributed across the memories of people who have been burned, and writing it down once converts it into something the next team inherits.

The decision log, including the noes. What you bought, what you declined, and why. Two paragraphs each. Kept somewhere findable by someone who does not know it exists.

Four documents, one owner, no budget. That is the whole apparatus.

The asymmetry you can actually close

This series opened with three asymmetries between the buyer and the seller. They know your process because they run it constantly. They know the market because they are in it every day. And they can wait longer than your sponsor’s attention lasts.

Two of those are not available to you. You will never have their pricing data, and you cannot out-wait them on any individual deal.

The third one is entirely yours. Their advantage is repetition, and repetition is a thing you can build. Not by buying more often, which would be a strange goal, but by making each purchase leave something behind that the next one starts from.

An organization on its fourth evaluation with none of that is on its first evaluation, four times. An organization that kept those four documents is genuinely, measurably better at this than it was, and it closed the only gap it was ever able to close.

You will still get some wrong

I have done most of what is in these fifteen lessons, and I have also skipped most of it, usually under time pressure, and some of those purchases were fine. Process is not a guarantee and anyone selling it as one is selling something.

What changes is the character of the mistakes. Without any of this, you make the same three errors repeatedly: no defined no, no measured baseline, no idea who could stop you. With it, you make new and more interesting errors, about things that were genuinely hard to see.

That is a reasonable definition of getting better at something, and it is the realistic version of what this series is offering. Not better outcomes every time. Failures that at least have not happened before.

One thing to do differently

When this evaluation closes, before the group disbands, spend two hours writing down what you would tell the person who runs the next one.

What you would ask earlier. Who you would have gone to see in week one. What the vendor said that turned out to matter. What you declined and why.

Then put it somewhere they will find it without knowing to look. That document is the capability. Everything else in this series is just a description of what tends to be written in it.

In Practice: The Other Side of the Table | Lesson 14: Renewal Is the Real Negotiation

You spent four months negotiating the first contract and you will spend about three weeks on the renewal. The renewal is the one that is worth more.

Six weeks and no alternative

Composite from a few multi-country service operations: a hotel group, guest service and revenue operations across several markets.

The original negotiation had been done properly. Competitive process, three bidders, a discount procurement was rightly pleased with.

The renewal arrived with, illustratively, a forty percent uplift, justified by usage growth that was entirely real. Procurement had six weeks. There was no second quote, because nobody had spoken to the runner-up in two years. Nobody knew what leaving would cost, because nobody had ever worked it out. And the vendor could see the usage data, because the buyer had been sending it to them monthly in the service review.

They negotiated it down somewhat and paid most of it. Which was, given where they were standing in week one of those six weeks, about the best available outcome.

The positions have swapped

At signature you had real alternatives and no dependency, and the vendor had better information about the market than you did.

At renewal that reverses. You now have excellent information about what the product is worth to you, which is a genuine advantage and the only one you gain. They have precise knowledge of how embedded you are, often from data you gave them voluntarily and correctly. And your alternatives exist in theory, in a slide, unrefreshed since the original evaluation.

The uncomfortable part is that most of your room to move exists in month four and is gone by month ten, which is roughly when anyone starts thinking about it.

Four things done early

Diarize it at signature. Not the renewal date. Renewal minus six months, with a named owner who is not the person most invested in the relationship going smoothly.

Keep the measurement running. The baseline you established before the vendors arrived, and the benchmark you wrote into the contract. Evidence of value is the only argument at renewal that is not a threat, and it is the only one that works when you have no credible way to leave.

Keep one alternative warm. An hour a year with the runner-up. It costs almost nothing, it makes a second quote obtainable in days rather than months, and it keeps you current on what the category can do now, which after two years may be considerably more than what you bought.

Refresh the switching cost. A number, from the leave test, updated once a year. Without it you cannot tell whether a proposed increase is worth accepting, and you will end up deciding on how the number feels.

Ask for things that are not price

Renewal is the best moment in the whole relationship to collect the terms you did not get the first time, because the vendor wants a clean renewal and non-price terms cost them less than a discount.

  • The benchmark clause, the deprecation notice and the export terms you traded away in round one
  • A commitment re-baselined to your actual usage rather than the forecast you got wrong
  • Price protection for the following term, so you are not doing this again from the same position
  • A unit price that reflects the fact that their input costs have almost certainly fallen since you signed

That last one is the two-line model from the cost lesson, and this is the moment it pays. If you tracked what you would be paying at something nearer list, you can tell whether the uplift is a correction of an introductory price or an increase on top of one.

The pressure from your own side

Worth naming, because it catches people who handle the vendor side well.

Your operating team does not want disruption. Your sponsor has moved on to something else and does not want this back on their desk. And somebody will say, accurately, that there is no capacity for a re-evaluation this quarter.

All of that is true and none of it is an argument about price. It is the condition the renewal number was set against, and a vendor with a decent account team knows the state of your capacity roughly as well as you do. The six months of preparation exist precisely so that the renewal does not require a re-evaluation.

One thing to do differently

On the day you sign, put an entry in the calendar at renewal minus six months, with a named owner and three items: pull the measurement, call the runner-up, refresh the switching cost.

Three hours of work, done once, eighteen months before it is needed. It is the highest-return thing in this entire series and it is skipped almost universally, because on signature day the renewal feels like somebody else’s problem and by the time it is yours the useful window has closed.

In Practice: The Other Side of the Table | Lesson 13: Exit Rights and the Leave Test

Somebody in the approval meeting will say that you can always switch later. It is usually said to close down a risk discussion, it usually works, and in most organizations it has never once been tested.

Eleven months to leave

Composite from a few operations-heavy deployments: a rail and freight operator, maintenance scheduling and incident triage across a large asset base.

Two years in they decided to move, for reasons that were mostly commercial and entirely reasonable. Illustratively, the migration took about eleven months and consumed roughly three times the annual license in internal effort.

The vendor did not obstruct. The data came out in a week, in a clean format, exactly as the contract said it would.

What did not come out was two years of accumulated fit. The thresholds somebody had tuned in month four. The exception rules that had grown to cover the seventeen situations the original design had not anticipated. The routing logic that assumed one product’s way of scoring severity. Most of it lived in a configuration interface, none of it was documented, and the two people who understood why it was set that way had both moved on.

The data was portable. The knowledge was not, and nobody had ever counted it as an asset.

What actually holds you

Roughly in order of how hard each is to move, which is close to the reverse of how much attention each gets in a contract negotiation.

  • Tuning and configuration. Thresholds, prompts, rules, weightings. Lives in a screen, rarely documented, exportable in no meaningful sense even when the vendor cooperates fully.
  • The shape of the work. Your process quietly reorganized itself around the product’s assumptions. Undoing that is a change program, not a migration.
  • Derived artifacts. The evaluation set, the corrections your users made over two years, the labeled examples. Often the most valuable thing you created during the contract, and often not clearly established as yours.
  • Integration surface. Connectors, identity pattern, reporting. Real work, but at least it is visible work that somebody can estimate.
  • Raw data. The thing everybody negotiates hardest for, and the least of the five.

The leave test

Before signature, write one page answering four things.

What we would get back, in what format, in how long. What we would have to rebuild rather than move. Who understands the configuration and whether that understanding exists anywhere outside their head. And what we would run on while we migrated, because the answer is rarely nothing.

Then do the version almost nobody does. During the pilot, ask for a full export of your data and your configuration, and time it.

A vendor who produces a clean export in a week is telling you something true about how they think about customers. A vendor who needs a professional services engagement to answer the request is telling you something else, equally true, and much better learned in week six of a pilot than in month twenty-six of a contract.

It also has a useful side effect. Asking during the pilot is a neutral, technical request. Asking during contract negotiation reads as distrust and gets handled by a different department.

What to write down

Four terms, in decreasing order of how easily you will get them: export format and timeframe stated explicitly rather than on request; ownership of derived artifacts established as yours; configuration included in what is exportable; and a transition assistance period at rates agreed now rather than at whatever they quote when you are leaving.

The third is the one that gets waved through as a technicality and is the one the rail operator would have wanted most.

Some of this is fine

Worth saying plainly: a product you cannot easily leave is often a product that has become genuinely useful. Deep fit and high switching cost are the same phenomenon viewed from two directions, and an organization that has embedded a tool into how work happens has usually got the value it paid for.

The goal is not zero switching cost. That would mean you never adopted anything properly.

The goal is knowing the number. A buyer who can say switching would take us about nine months and cost roughly this much can make a rational decision at renewal, including the decision to stay and pay more. A buyer who has never worked it out negotiates as though they could leave in the spring, then discovers in week two that they cannot, having already made a suggestion they now have to walk back.

One thing to do differently

During the pilot, request a full export of your data and your configuration, and time how long it takes to arrive.

You will learn more about your future switching cost from that one request than from the entire exit rights section of the contract, and you will learn it while you still have three vendors and no commitment.

In Practice: The Other Side of the Table | Lesson 12: Contracting for Something That Changes Under You

You are about to sign a contract for a product that will not be the same product in six months, using a template written for software that shipped twice a year and changed when you accepted an upgrade.

Nothing in this lesson is legal advice. It is the commercial shape to hand to the person who does give it, because your legal team will draft what you ask for and they are not going to know to ask for these on their own.

The change that breached nothing

Composite from a few archive-scale evaluations: a media and publishing group, automated content classification and rights checking across a large back catalogue.

They deployed, then spent about three months tuning their thresholds until the review queue was a size two editors could handle. Good outcome, sensibly done.

Then an update on the provider side changed how the underlying model behaved. Nothing broke. Illustratively, precision moved a few points one way and recall a few points the other, which is the sort of change that is invisible in aggregate and very visible in a queue. Their carefully tuned thresholds were now set for a system that no longer existed.

They found out roughly six weeks later, when an editor mentioned the queue felt heavier. No term of the contract had been breached. No term of the contract covered it. Their own vendor had not been told in advance either.

Five terms worth asking for

Notice of material change to the underlying model. Vendors resist the word material, and fairly, because it is vague. Offer a test instead of an adjective: any change of model family or version, plus anything they announce publicly as a behavior change, plus anything they themselves detect as moving output distribution.

Deprecation notice with a floor. How long you get before a version you depend on is withdrawn. Ask for a stated number of months. Ask separately for the right to remain on a prior version during a validation window, which is worth a great deal in regulated work and very little elsewhere, so do not spend negotiating capital on it if you do not need it.

A benchmark written into the agreement. Almost nobody does this and it is the most useful of the five. Agree a small set of your own held-out cases and an acceptance threshold at signature. Re-run it quarterly. If the result falls below the threshold, you have a defined remedy, even if the only remedy you can win is a termination right.

Without it, the sentence available to you in month eight is that it seems worse than it used to be, which is an opinion, and you will be having an argument rather than exercising a right.

Your data out of training, in terms that survive a product change. Most vendors will give you this readily. What varies is whether it covers the derived artifacts from the previous lesson, and whether it binds their providers as well as them.

What happens if their provider relationship ends. Many vendors in this category are built on capability they license. Ask what happens if that arrangement changes, whether they can serve you on an alternative, and how long that would take. The answer tells you how much of the product is actually theirs, which is useful well beyond the contract.

What you will actually get

Not all five. Possibly two. So rank them before you walk in, because the order in which you concede is decided in the room otherwise, and in the room the unfamiliar clause always goes first.

For most buyers the benchmark and the deprecation notice are worth more than the price concession you would trade to keep them. Procurement will not see it that way, because a discount appears in the savings number and a benchmark clause appears nowhere. That is a reporting artifact rather than a judgment, and it is worth naming out loud before the negotiation rather than after.

One tactical note. Ask for these early, in the first commercial conversation, framed as how we work rather than as demands. A clause introduced in week two is a requirement. The same clause introduced in week nine, after the shortlist is public and your sponsor has told the board, is a problem you are creating for a deal everyone now expects to close.

The list you already have

If you have been keeping the running list from your reference calls, the answers to what would you put in the contract if you were signing again, this is where it earns its keep.

Those answers tend to be narrow and specific in a way that generic templates are not. A named support contact. A cap on the annual uplift. The right to add users mid-term at the original rate. Notice on a change to how consumption is counted. Every one of them came from somebody who learned it at their own expense, and there is no cheaper source of contract language anywhere in this process.

One thing to do differently

Put a benchmark in the contract. Twenty of your own cases, an agreed threshold, re-run quarterly, with a stated consequence if it fails.

It takes an afternoon to assemble and it converts the most likely failure in this category, quiet degradation that nobody can prove, from a disagreement into a clause.

In Practice: The Other Side of the Table | Lesson 11: Diligence That Is Not a Questionnaire

You have a security questionnaire. It runs to several hundred rows, it was written for hosted software, and the vendor will complete it in two days using answers they have given ninety times before.

What comes back tells you the vendor can complete a questionnaire. That is a real signal about organizational maturity and it is not the same thing as having done diligence.

Two weeks to answer a simple question

Composite from a few evaluations in devolved organizations: a research university group, several faculties, federated IT, a purchase covering research administration and grant documentation.

The questionnaire came back complete. Certifications attached, everything in order, no red flags. Procurement signed it off in a week, which was fast and felt like a win.

About nine months later a faculty raised a query about where a particular category of data had been processed during the previous quarter. Nothing had gone wrong. Nobody had been breached. But a subprocessor in the chain had changed, the notice had gone to a shared procurement mailbox that nobody monitored, and it took roughly two weeks to assemble an answer to a question that should have taken an afternoon.

The questionnaire had asked whether the vendor used subprocessors. It had not asked how the university would find out when the list changed, and that is the question that mattered.

Why the form misses

Four structural reasons, none of which involve anyone cutting corners.

It asks about controls in general rather than about your deployment in particular. It is answered by a compliance function rather than by the people who operate the system. It captures a point in time, and the product you are buying changes monthly. And it was written for a category where the software you license is the software you run, which is not the arrangement here.

That last one matters most. Most products in this category sit on capability the vendor does not own, served from infrastructure they do not run, and the questionnaire has no row for that.

Five questions instead

Run these as a conversation, forty-five minutes, with the vendor’s engineering lead in the room rather than their compliance team. Insist on that. The whole value is in the follow-up questions, and a compliance function cannot answer follow-ups.

One. Where does our data go, every hop, and how will we be told when that list changes?

Not whether they use subprocessors. The list, the notification mechanism, the notice period, and which address it goes to. The mechanism is the part everyone forgets and the part that failed at the university.

Two. What of ours is retained, for how long, and what is derived from it that survives deletion?

Deleting source data does not necessarily remove what was built from it. Indexes, embeddings, caches, logs, evaluation sets, corrections your users made. This is the question with the widest range of answers in the industry right now, which is exactly why it is worth asking.

Three. Who at your company can see our data, in what circumstances, and can we see that access?

Support debugging is the usual route and it is entirely legitimate. You want to know the path exists, what triggers it, whether it requires customer approval, and whether the log is visible to you or only to them.

Four. What changed in the last twelve months that you were not obliged to tell customers about?

This surfaces the entire class of change that sits outside contractual notice provisions. It is also an excellent question for a reference call, where you will sometimes get a more complete answer.

Five. Show me your last incident and the write-up.

Not whether they have had incidents. Every operation running at scale has. A vendor who produces a post-incident review, even heavily redacted, is showing you how they operate more convincingly than any certificate. A vendor who says they have not had one is either very new or not counting the same things you would.

Keep the form

None of this is an argument for skipping the questionnaire. Send it. Your own audit function needs it, a vendor who cannot complete one is telling you something real about their maturity, and comparing completed forms across three vendors does occasionally surface a genuine difference.

The mistake is treating a completed form as a finished investigation. The form is the floor. The forty-five minute conversation is the diligence, and it costs less than the week your team will spend reconciling three completed spreadsheets.

One thing to do differently

Ask question two in writing, and keep the written answer.

What is derived from our data, and what survives deletion. It is the question where answers vary most between vendors, it is the one your data protection colleagues will come back to in year two, and a written answer given during evaluation is worth considerably more than a verbal reassurance you half remember.

In Practice: The Other Side of the Table | Lesson 10: Seats, Tokens, Outcomes

You’re being offered a price in one of three shapes. The shape will do more to determine what you pay over three years than the number attached to it, and it’s the part most evaluations spend the least time on.

What each shape does to you

Per seatConsumptionOutcome
You pay forPeople with accessVolume usedResults delivered
BudgetsCleanlyBadlyUnpredictably
Bad incentive it createsRestrict who gets accessDiscourage useArgue about definitions
Who carries the varianceThe vendorYouContested
Right whenValue is per person and adoption will be broadUsage is concentrated in a few heavy usersThe outcome is already measured in a system you both trust

The row worth sitting with is the third one. Every pricing shape creates an incentive that works against the reason you bought the thing.

Per-seat pricing makes every additional user a cost, so organizations ration access, and a tool that only the enthusiasts can reach never changes how the work is done. Consumption pricing makes every use a cost, so managers tell their teams to be thoughtful about it, which is a polite instruction to use it less. Outcome pricing makes the definition of the outcome the thing worth arguing about, and you will argue about it, in month seven, with someone whose commission depends on the answer.

You can’t escape this. You can pick which version of it you’d rather manage.

Paying for a curve that didn’t happen

Composite from a few regulated-documentation evaluations: a pharmaceutical company, medical writing and regulatory submissions, a few hundred people in scope across two sites.

What they signed was the shape most of these contracts settle into: a platform fee, plus consumption, with a committed annual minimum in exchange for a better unit rate. Illustratively, the commitment was set at roughly the volume they expected in month nine of year one, extended across the year.

They reached about forty percent of that in year one. They paid for all of it.

The commitment wasn’t a trick and the vendor had given real value for it. The error was upstream: the adoption forecast came from pilot volunteers, who had used the tool enthusiastically for ten weeks, and the curve was drawn as though four hundred medical writers would behave like the thirty who’d signed up.

Adoption in a large regulated function is slow for reasons that have nothing to do with the software. Validation, standard operating procedures, someone has to update a template, a therapeutic area lead is on leave. None of that was in the model.

The clauses that make a commitment survivable

Committed minimums are not inherently bad. They’re how you get a decent unit rate, and refusing all commitment usually means paying a premium for flexibility you may not need. Four terms make them liveable:

  • Rollover of unused commitment into the following period. The most commonly granted of the four and the most commonly not requested.
  • A ceiling on unit price, not only a floor on volume. If you’re guaranteeing their revenue, they can guarantee your rate.
  • A re-baseline right at twelve months, resetting the commitment to actual usage. Harder to get, worth asking for, and the reaction tells you how confident they are in your adoption.
  • Protection against changes underneath. If your consumption is measured in units the vendor defines, and the definition changes because the technology changed, your effective price moves without anyone renegotiating.

That fourth one is specific to this category and it’s newer than most procurement templates. Consumption units in AI products are not stable objects. A change in how the product processes a request can change your unit count materially, in either direction, without a single term of the contract being amended.

The thing almost nobody negotiates

Ask what happens to your price when their costs fall.

In a category where the underlying cost per unit of capability has been falling consistently, a three-year fixed price is a bet you’re making against the trend, and you’re making it silently. An annual benchmark review, or a clause that passes through some share of a reduction in underlying cost, is unusual but not absurd to ask for.

Most vendors will decline. The way they decline is informative. A vendor who explains that their own input costs are contracted and can’t be passed through is telling you something specific and probably true. A vendor who treats the question as inappropriate is telling you how the renewal conversation will go.

Which to insist on

If the value is genuinely per person and you intend broad rollout, per seat is usually right, and negotiate the right to add users mid-term at the original rate. That last clause is cheap to get at signature and expensive to get later.

If usage will be concentrated in a small group doing heavy work, consumption is right, because per-seat pricing across a large population to serve twenty real users is how you end up with an unflattering cost per active user in the benefits review.

Outcome pricing is right when the outcome is already being measured, by a system you both already trust, that neither party controls. In practice that condition holds rarely. When someone offers it and the condition doesn’t hold, what’s being offered is a future dispute with a discount attached.

One thing to do differently

Model whatever shape you’re offered at three adoption levels: the rate your pilot volunteers achieved, half of your plan, and double it.

The first is the fantasy, the second is the likely case, the third is what happens if it works better than expected, which is a scenario people rarely cost and which has bankrupted more than one committed-volume contract from the other direction. Sign the shape that’s survivable in all three.

In Practice: The Other Side of the Table | Lesson 9: Reading a Cost Model From the Outside

The quote in front of you is one number with a discount attached to it. Underneath that number is a cost structure that determines whether the price is stable, and nobody is going to show it to you.

You can infer the shape anyway. Not the numbers, the shape, which is enough to tell you which parts of your price are durable and which parts are being subsidized until renewal.

A price that was never going to hold

Composite from a few knowledge-work evaluations: a professional services firm, a couple of thousand fee earners, document review and drafting support across several practice areas.

They negotiated hard and negotiated well. Illustratively, something like fifty-five percent off list, three-year term, annual true-up on volume. Procurement was pleased and had every right to be.

Adoption then grew, which was the entire point of the purchase. By year three the firm was running several times the volume it had modeled. At renewal the price roughly tripled, and the argument for the increase was the firm’s own success.

Nobody was deceived here. The original number had been an acquisition price, priced to win a logo in a competitive segment, and the vendor had been reasonably transparent that volume growth would change things. What nobody on the buying side had done was ask which parts of the price moved with usage and which parts didn’t. They’d negotiated the total, which is the one number that told them least.

Four layers under the number

Every price in this category sits on the same four layers. You can’t see the values. You can reason about the behavior.

Compute. Moves with your usage. Falls over time as models get more efficient, but never to zero, and it doesn’t fall on a schedule anyone can commit to. This is the layer that makes a flat price unstable when your volume grows.

The product around it. Their engineering, interface, and everything that makes the underlying capability usable. Fixed cost, spread across all customers, so the per-customer cost falls as they grow. Cheap for them to be generous with, which is why feature access is usually the easiest concession to win.

Cost to serve you specifically. Implementation, your custom connector, your security review, the quarterly business review, the named support contact. Large, lumpy, and mostly independent of how much you use the product. This is why a small contract often has worse unit economics than it looks and why vendors push hard on term length for smaller accounts.

Acquisition. Sales, marketing, and the discount itself. Recovered across the expected life of the account, which is precisely why the term matters more to them than the annual figure. A three-year commitment at a deep discount can be a better deal for them than one year at list.

Reading the concessions

What a vendor gives away easily and what they defend tells you where their cost sits. This is the most reliable instrument you have, because it works on behavior rather than on claims.

What they doWhat it suggests
Deep discount for a longer termAcquisition cost is being amortized. Expect a reprice at renewal.
Feature tiers collapse easilyThe product layer is cheap to them. Ask for more of it.
Accepts a cap on unit price, resists a cap on volumeTheir per-unit cost is falling and they expect to grow into it. Good sign for you.
Accepts a volume cap, resists a price capThe reverse. Your unit cost is exposed.
Resists both, offers a bigger headline discountThe discount is the concession and it expires. Read the renewal terms first.
Implementation fee is non-negotiableCost to serve is real and they’re not absorbing it. Usually honest, sometimes a partner margin.

None of these are certainties. They’re priors, and they’re much better than the prior you get from the proposal document, which is written to be read in one direction.

The question that gets you the shape

There’s one question that does most of the work, and it isn’t hostile, which is why it gets answered:

Which parts of this price move if our usage doubles, and which parts don’t?

A vendor with a considered pricing model answers it in a couple of minutes and often quite openly, because it’s a question about structure rather than about margin. A vendor who can’t answer it is either not close to their own economics or is hoping you won’t model it, and both are worth knowing.

Then take it further and ask for the same quote at half your projected volume and at double it. Three points on a line show you the shape: whether your unit price improves with scale, stays flat, or quietly worsens past a threshold you hadn’t spotted.

I’ve had a vendor decline to produce the double-volume quote on the grounds that it was hypothetical. It was hypothetical. It was also the volume in their own business case for us, which they had presented the previous month.

Where the subsidy lives

A price below the vendor’s cost to serve you isn’t a victory. It’s a repricing with a date on it.

That doesn’t make it a bad deal. Being an early customer in a category that’s subsidizing growth is often genuinely good value, and the subsidy is real money you get to keep for the term. The mistake is treating the introductory price as the price and building a three-year business case on it.

So model two lines. What you’ll pay under the contract, and what you’d pay at something closer to list with your projected volume. If the business case only works on the first line, you’ve bought a discount rather than a capability, and you’ll find that out at renewal when your usage is embedded and your alternatives have narrowed.

One thing to do differently

Ask for the same quote at half your projected volume and at double it, before you negotiate the headline number.

It takes them a day. It tells you more about what you’re signing than any amount of discussion about the discount, and it moves the conversation from a single number you can argue about to a structure you can plan against.

In Practice: The Other Side of the Table | Lesson 8: Reference Calls That Get Real Answers

You will be given three references. All three are current customers, all three were asked whether they’d take the call before their names reached you, and all three have been told roughly what you’re evaluating.

That process is designed to produce a positive signal and it works. The interesting question is what you can get out of it anyway, and where the rest of your sample comes from.

Three happy customers, one region

Composite from a few infrastructure-heavy evaluations: a telecommunications operator, network operations and field dispatch, a workforce spread across three countries and several timezones.

They took the three references offered. All three were positive and none were dishonest. The buyer signed.

In month four they discovered that the vendor’s support model had real coverage in one region and a follow-the-sun arrangement everywhere else that mostly meant a ticket sat overnight. Their night shift was the shift that most needed it.

All three references had been in the well-covered region. Not chosen for that reason, as far as anyone could tell. That’s simply where the vendor’s happiest customers were, which is the same thing as saying that’s where their support was good.

The reference sample and the quality signal were the same variable, and nobody noticed because nobody asked what the three had in common.

Questions with factual answers

Do not ask a reference whether they’re happy. They were selected for being happy and they’ll confirm it accurately.

Ask questions with factual answers instead, ideally ones involving dates and artifacts, because those are hard to soften without lying and almost nobody wants to lie to a stranger on behalf of a vendor.

  • How long from signature to your first real user? A date minus a date. If it’s four times the vendor’s stated implementation time, that’s your number, not theirs.
  • What surprised you in month six? Month six is after the honeymoon and before the renewal, which makes it the most honest window in the relationship.
  • What did you end up building that you thought you were buying? Every implementation has at least one. The answer is your integration estimate.
  • Who on your team would have voted against this, and what was their argument? This gives a satisfied customer permission to voice the internal objection without owning it.
  • If you were signing again, what would you put in the contract?

That last question is the best in the set and I’d trade the other four for it. It hands you contract terms you don’t yet know to ask for, from someone who learned them at their own expense and has no reason not to share.

The answers tend to be specific and unglamorous. Notice periods on model changes. A cap on the annual uplift. A named support contact. The right to add users mid-term at the original rate. Write every one of them into your contracting checklist as it arrives.

The rest of the sample

Three selected references are a sample of one population: satisfied customers, in the segment the vendor serves best, at their current maturity. You need at least one point outside that.

Former customers. The hardest to find and the most informative. Your own network is the usual route. Peer groups and industry associations are better than they sound, because someone in the room has always left someone.

Organizations that evaluated and chose otherwise. Often easier to reach, and they’ll tell you exactly what tipped it, which is usually one specific thing rather than an overall judgment.

Old case studies. Take the vendor’s published customer stories from three years ago and check whether those organizations are still customers. It costs an hour. A published reference who has since moved on is worth a call.

Illustratively, on one evaluation I’d put the three given calls at around eight useful minutes each. One conversation with a company that had left the vendor eighteen months earlier produced more than all three combined, largely because they had nothing to be careful about.

Weighting the unhappy one

A churned customer is also a biased sample, and it’s worth saying so plainly rather than treating the negative view as the true one.

They left for a reason. The reason may be specific to their data, their sector, their sponsor leaving, or a botched implementation by a partner the vendor no longer works with. Any of those can be irrelevant to you.

What you’re after isn’t the negative account. It’s the variance. Three positive calls tell you the product works somewhere. A fourth call from outside that population tells you how wide the band is between the best case and a bad one, and the width of that band is the actual risk you’re taking.

One practical thing: ask each of the three given references what they have in common with the others. Sometimes they know. Same industry, same region, same implementation partner, all onboarded by the same person. If all three share a trait you don’t share, you’ve learned something the calls themselves were never going to tell you.

One thing to do differently

Ask every reference what they’d put in the contract if they were signing again, and keep a running list across all the calls.

By the fourth conversation you’ll have a set of clauses assembled from other people’s expensive lessons, which is a considerably better starting position than a template your legal team last updated for a hosting agreement.

In Practice: The Other Side of the Table | Lesson 7: The Demo Is a Product Too

The demo you’re watching was built, tested and rehearsed. It has its own backlog and its own maintainer. The person delivering it has given it more times than you have bought enterprise software in your career.

You are evaluating that artifact. You are not, yet, evaluating the product.

When every demo is excellent

Composite from a few casework evaluations: a public sector agency, permit and licensing, several hundred caseworkers, a document-heavy assessment process with a long backlog.

Three vendors demoed. All three were excellent. Genuinely, not sarcastically: polished, responsive, well-matched to the described problem. The evaluation team came out of three sessions with three sets of enthusiastic notes and no way to tell the products apart.

So they decided on price, which felt disciplined. Eleven months later the product was struggling with the agency’s actual document mix, a large share of which arrived as scans of forms that had been filled in by hand and then photocopied. None of the three demos had included a scan of anything.

Nobody misled anyone. The agency had never asked, and no vendor volunteers the input type their product handles worst.

What a demo selects for

A demo isn’t deceptive. It’s selective, and the selection happens along exactly the three axes that decide whether the thing works for you.

The path. It follows the route through the product where the product is strongest. Every product has one. Yours will not be it.

The data. Demo data is clean, representative, and often assembled specifically because the product handles it well. Your data has a decade of exceptions in it and four fields that mean different things depending on which system wrote them.

The scale. One user, a handful of documents, no concurrency. Latency at your volume is a different product experience and it never appears in a demo.

Underneath those sits the thing that matters most and is hardest to see: what it took to make it look like that. Setup effort is invisible in a demo by construction, and setup effort is most of your integration budget.

Three requests that break it

Three, not ten. A long list turns the session adversarial and you’ll get defensive answers to all of them. These three are enough, and you can make all three sound like curiosity rather than cross-examination, because that’s what they are.

One. Run it on our data, unprepared, now.

Not a curated sample sent two weeks ahead. A file pulled from your system that morning, exceptions included, in whatever state it’s actually in. Bring three.

You’re not testing accuracy. You’re testing what happens on input the vendor didn’t get to shape. If the honest answer is that this needs a preparation step first, you’ve just found part of your integration cost, in week three rather than month five. That’s a good outcome from the request, not a failed one.

Two. Show me it being wrong.

Ask them to produce a case where their system fails, and then to walk through how the user finds out.

This is the single most informative question in an evaluation. A vendor who has a failure case ready, describes it precisely, and can explain the detection path has thought hard about operating in the real world. A vendor who deflects has either not characterized their failure modes or has decided not to discuss them, and you want to know which before you’re a customer rather than after.

The follow-up matters as much: how does the user know. A system that’s wrong loudly is manageable. A system that’s wrong quietly, in a document somebody signs, is a different risk profile entirely and belongs in your business case rather than in a footnote.

Three. Who built this demo, and are they on our implementation?

Nearly always the answer is a solution engineer who is not. That’s not a scandal, it’s how the industry is staffed, and treating it as a gotcha wastes the question.

The value is in what comes next. Who does implement, how many of those people exist, how many concurrent implementations they’re running, and whether it’s the vendor’s own team or a partner. You’ll learn more about your delivery risk in that four-minute answer than in the entire capability section of the RFP response.

Who should be watching

Put two people in the room who will personally use the thing, and tell them beforehand that being difficult is their job today.

They will ask things you wouldn’t. A caseworker watching a demo of casework software notices in about ninety seconds that the screen assumes one applicant per file, and says so, while the evaluation team is still admiring the summarization. That observation is worth the entire session.

The reason this doesn’t happen by default is that demos get scheduled as executive sessions, and adding practitioners feels like it dilutes the audience. It’s the reverse. The executives can read the summary.

One thing to do differently

Ask for the failure case in the first demo rather than the third.

It changes the relationship immediately, in a way that’s mostly good. It signals you’re a buyer who’ll be dealing with the product in production, not a buyer being impressed. And it sorts your shortlist within ten minutes, because the vendors who’ve thought seriously about being wrong answer it well and are visibly pleased to be asked.

In Practice: The Other Side of the Table | Lesson 6: Pilot Theatre

You have agreed a pilot. Ten weeks, two vendors, a defined scope, a steering group, and nowhere in any of the documents is a sentence describing the result that would make you stop.

Pilots almost never fail. That is the tell. If your organization has run six of these and all six succeeded, the pilot is not measuring anything, it’s producing a decision that was already made and giving everyone something to cite.

The criteria that cannot produce a no

Composite from a few clinical deployments: a hospital group, several sites, a discharge documentation problem. The pilot ran ten weeks with forty clinicians.

The success criteria, as written, were to demonstrate feasibility and gather user feedback. Both of those are guaranteed before anyone starts. Feasibility was demonstrated, in that the system produced discharge summaries. Feedback was gathered, in that people had opinions. Illustratively, sixty-eight percent of participating clinicians said they would use it again, and that number became the headline in the steering pack.

What nobody measured was how heavily the drafts were edited, and whether editing took longer than writing from scratch on the complex cases. When someone finally looked, roughly the hardest fifteen percent of discharges took longer with the tool than without it. Those are also the ones where errors matter most.

The product was fine. It probably was worth buying, scoped differently. But the pilot had been built to produce a yes, and it produced one, and the scoping conversation that should have happened at week ten happened at month fourteen instead.

Four reasons your pilot will pass

None of these involve anyone behaving badly. They’re structural, which is what makes them reliable.

The vendor is staffing it. There’s a solution engineer in your channel answering within the hour. That is not the support model you’ll have in production, and it is quietly fixing problems you never find out existed.

Volunteers, not conscripts. The forty people who signed up are not a sample of the four hundred who’ll be told to use it in March. They’re the ones who like new tools. Their enthusiasm is real and it does not generalize.

The easy slice. Ten weeks is not long enough to assemble messy data, so pilots run on the clean subset. The clean subset is not where your cost is.

Accumulated commitment. By week ten the organization has spent four months, three workshops and a steering group’s attention. Saying no at that point costs someone social capital, and everyone in the room can feel it.

Building one that can fail

A pilot that can fail is not a hostile pilot. It’s the only kind that tells you something you didn’t already believe. Six things make the difference:

  • One number, measured the same way before and during. Not a dashboard. One.
  • A written kill criterion, agreed by the sponsor before the vendor is told the pilot is happening. If it’s written after, it will be written to be passable.
  • Reluctant users in the population. Pick five people who didn’t volunteer and ask them to take part anyway. Their experience is the one that predicts month nine.
  • The hard cases in scope. Name the ugliest ten percent explicitly and require them to be included, because that’s where the business case lives or dies.
  • A declared support level. Ask what this looks like without a solution engineer in the room, then run at least the last three weeks that way.
  • A named person who can stop it. Not a committee. One person, told in advance that stopping is an acceptable outcome and will not be held against them.

That last one costs nothing and is skipped almost every time. A committee cannot stop a pilot. Committees produce continuations with caveats.

The exit nobody writes down

Before the pilot starts, settle three things in writing. They take a paragraph each and they’re painful to negotiate afterward.

What happens to your data when the pilot ends, including anything derived from it. Whether outputs produced during the pilot can be used if you don’t proceed. And whether the pilot fee, if there is one, credits against a contract, because a pilot fee that only credits on signature is a small, entirely legal incentive pointed at your own decision.

The outcome your process cannot record

Sometimes the honest pilot result is that the thing works and isn’t worth the change cost. It does what it says, the number moves a little, and getting four hundred people to alter how they work costs more than the gain.

That is a successful pilot. It is also a failed purchase, and most governance processes have no box for it. The options are proceed, or proceed with conditions, or defer, and defer is where good projects go to be quietly embarrassing.

If you can add one box to your template this year, add that one. Works, not worth it. I’ve watched a room spend forty minutes trying to phrase that conclusion in a way that wouldn’t look like failure, and land on deferring pending further review, which is how a clear finding turns into eight months of nothing.

One thing to do differently

Before the pilot starts, write the single sentence that would end it. Then get the sponsor to sign that sentence, not the pilot plan.

The signature is the part that matters. It converts stopping from an act of individual courage in week ten into a thing that was agreed in week one, by someone senior, when it was still cheap to agree to.