The Peacock arrived for the second planning cycle with a display of quite exceptional magnificence.
There were slides. There was a total cost of ownership model with a five year horizon and a curve on it that started high and ended gratifyingly low. There was a comparison against what the jungle was currently paying, and the comparison was not close. Somebody in the room said the words game changer without any apparent irony.
The Fox, who had learned a great deal over the preceding two acts, asked a reasonable question about the assumptions.
The Crocodile, who had sat through this presentation four times in his career under four different names and two different technology paradigms, opened one eye and asked the only question that actually decides it.
“What utilization is that curve assuming?”
There was a pause of a length that told everyone in the room what the answer was going to be before the Peacock said it.
Peacocks have enormous displays and modest payloads. This is not dishonesty, it is what the animal is for. The tail is the product. Somebody in the room has to be the one who asks about the bird underneath it, and in most organizations there is no such animal, which is why these decisions go wrong.
Two shapes, not two products
The frontier versus local argument gets conducted as though it were a choice between two products with different features. It is not. It is a choice between two cost shapes, and almost everything else people argue about is downstream of that.
API-served inference is pure variable cost. Near zero fixed. You pay for what you use, forever, and the line is straight. It never gets cheaper through your own effort, only through the provider’s price changes, and it never costs you anything when you are not using it.
Self-hosted inference is high fixed cost with low marginal cost. You pay a large amount whether or not anybody uses it, and then very little per unit on top. The unit economics are terrible at low volume and excellent at high volume, and the entire question is where those two lines cross.
That is the whole thing. Everything else is a variation on it. And it means the honest form of the question is not which is better, it is what volume gets you past the crossover and how confident you are that you will be there.
The utilization trap
Here is where most self-hosted business cases fall over, and it is always the same place.
Fixed-cost economics are quoted at high utilization, because that is the only way the numbers look good, and that is a perfectly reasonable way to describe theoretical capability. They are then delivered at whatever utilization your actual workload produces, which is a completely different number.
Enterprise inference demand is not smooth. It follows the working day, which means it is close to nothing for a substantial portion of every twenty four hours. It follows the working week. It follows the business cycle, the month end, the campaign calendar. It has peaks that determine how much capacity you must provision, and troughs that determine how much of that capacity earns nothing.
You provision for the peak. You are billed for the peak. You utilize somewhere considerably below the average, once you account for the redundancy you need in order to survive the peak going wrong.
An honest utilization assumption for a first-generation enterprise deployment is dramatically lower than the one in the presentation, and the gap between those two numbers is usually the entire business case.
What is missing from the fixed side
The second failure is that the fixed side of the model is systematically incomplete. Four things that belong there and usually are not.
- The people. Self-hosting is an operational commitment, not a purchase. Somebody keeps it running, patches it, handles the incidents, manages capacity. That is a standing team, and in most models it appears as a footnote or not at all.
- Redundancy. A single deployment is a single point of failure for something you have just made business critical. The resilient version costs meaningfully more than the demonstration version.
- Evaluation and update cycles. Every model change you make yourself requires you to re-verify behavior across everything that depends on it. On the API side somebody else absorbs a version of this cost. On your side it is yours, and it recurs.
- Falling behind. This one has no line item and it is real. Capability moves. A fixed deployment holds still. The gap between what you have and what is available compounds, and at some point it becomes a cost in the form of work your organization cannot do.
The three questions
What I would ask, in this order, of any proposal in either direction.
What is the break-even volume? A specific number of units per period at which the two curves cross, with the fully loaded fixed side included. If nobody can produce this number, the analysis has not been done, whatever else has been produced.
How confident are you in the volume forecast? Not what is the forecast. How confident. Because the forecast is doing all the work and it is usually the least examined input in the model. Ask when it was made, by whom, and what it is based on.
What does this look like at half that volume? This is the question that decides it. Almost every self-hosting case survives the first two questions and fails the third, because the fixed side does not move and the unit economics collapse. If the answer at half volume is still acceptable, you have a genuinely robust case. If it is catastrophic, you are making a bet on a forecast rather than an architecture decision.
What optionality is worth
The thing that is almost never priced, and which I think matters more than anything else on this list.
Variable cost buys you the right to change your mind. Fixed cost buys you a lower unit price and a commitment to a capability level that will be surpassed, on a timeline you do not control, while you are still depreciating the thing.
In a slow-moving field that trade is usually worth taking, which is why organizations have made it comfortably for decades with every other kind of infrastructure. In a field where capability and price are both moving quickly, the option to move has genuine value, and the discount you are being offered has to exceed it.
The framing I would use in the room: what would you pay, today, for the ability to shift this entire workload somewhere else in six months with no stranded cost? Whatever that number is, it is the hurdle the discount has to clear. Nobody ever calculates it, and it is frequently larger than the saving.
The answer is usually both
For most enterprises the correct answer is not one or the other, and treating it as a single organizational decision is how it goes wrong.
Work that is high volume, narrow, well specified, stable in its requirements, latency sensitive, or subject to a constraint on where it can run, suits a fixed-cost local deployment. The volume is there, the capability requirement is not moving, and the crossover is genuinely favorable.
Work that is low volume, broad, hard, exploratory, or benefits from the newest capability suits variable-cost API-served inference. You are buying capability and optionality rather than unit price, and both are worth the premium.
The organizations that get this right decide per workload and build the routing layer that makes the decision reversible. That routing layer is the next lesson, and it is the mechanism that turns this from a one-time architectural bet into something you can adjust quarterly.
Three ways this goes wrong
The full utilization business case. Economics quoted at a utilization the workload will never reach, delivered at half of it, defended for two years by people who cannot admit the input was wrong.
The forgotten team. Hardware and capacity modeled carefully, the standing operational function required to run it modeled not at all. This single omission has sunk more self-hosting cases than every other factor combined.
Ideological procurement. The architecture chosen first, for reasons of principle or preference, and the model constructed backwards to support it. This happens in both directions and it is equally expensive either way.
The Field Kit
Concrete things to do this week.
If you sit in the Crow’s chair, ask the three questions in order and do not accept a break-even volume without a stated confidence in the forecast underneath it. Then ask for the version at half volume before you approve anything.
If you sit in the Crocodile’s chair, model the fully loaded fixed side including people, redundancy, evaluation cycles and refresh. You are the only animal in the building who can produce that number honestly, and it is the number that decides it.
If you sit in the Mandrill’s chair, be honest with yourself about whether your volume forecast is a forecast or an aspiration. The architecture decision is far more expensive to reverse than the forecast is to correct.
For everyone: decide per workload, never once for the organization. A single global answer to this question is always wrong somewhere, and it is usually wrong somewhere expensive.
Jungle Lesson 11
This is not a choice between two products, it is a choice between two cost curves, and the only questions are where they cross and how much you trust the volume that gets you there. Variable cost is the price you pay for the right to change your mind, and in a field moving this fast that right is worth more than most business cases give it credit for.
Next time: the Beaver builds something over a weekend that takes a large bite out of the bill, and nobody notices for a month because the outputs did not change at all. The Fox is mildly annoyed that the biggest win of the quarter came from plumbing rather than strategy. Lesson 12 is about the routing layer, and why the cheapest sufficient answer beats the smartest one.
If a proposal in either direction cannot tell you what it looks like at half the assumed volume, it is not an architecture decision. It is a forecast with an architecture attached, and forecasts are much cheaper to be wrong about.