Writing

The Cost Curve Nobody Models

Routing between local and frontier models is an architecture decision, not a procurement one

Image Two lanes

Midjourney prompt
painterly editorial illustration, two parallel channels of light diverging at a switching junction, one narrow and warm, one broad and cool, deep navy surroundings, warm amber in the near channel, soft teal in the far one, brass switching gear and worn rails, quiet engineering order, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
Two lanes The question is not which model. It is which work goes down which lane.

Enterprise AI business cases almost always take the same shape. Pick a model, estimate a volume, multiply by a published price, compare against the labor it replaces. The number comes out favorable, the program gets funded, and the actual bill arrives shaped differently.

The reason is not that the estimate was careless. It is that the model priced a single workload against a single provider, when the thing being built is a routing problem with several cost variables that move independently.

The decision is not which model to buy. It is which work goes down which lane, who owns that choice, and whether the architecture lets you change your mind later without a rewrite.

What the spreadsheet usually gets wrong

Price per token is the most visible variable and close to the least important. Four others move the total more, and three of them grow with success rather than staying flat.

Volume grows with adoption, non-linearly. A pilot serving a team is not a tenth of the cost of one serving the department, because the department finds uses the pilot team never proposed. This is the good outcome and it is the one that breaks the forecast.

Agentic work multiplies calls per task. A single user request is rarely a single call. It is retrieval, then reasoning, then a tool call, then a check, sometimes a retry after a failure. The multiplier between "requests" and "calls" is set by the architecture, is easy to change accidentally, and is frequently absent from the model entirely. A modest change in retry policy can move the bill more than switching providers.

Context size is a design choice masquerading as a requirement. Most systems send far more context than the task needs, because sending everything is easier than deciding what matters. That cost is paid on every call forever, and it is one of the few places where a week of engineering has a permanent multiplicative effect.

Figure The variables that actually move the number

Graph prompt
Draw a clean editorial ranking diagram titled What actually moves the cost of an enterprise AI workload, ordered by impact, largest at the top. Row one, Call volume and how it grows with adoption. Row two, Retries, retrieval, and multi step reasoning, the hidden multiplier on every call. Row three, Context size per call, driven by how much you send, not how much you need. Row four, Human review time on the output. Row five, Data movement and egress. Row six, Price per token. Draw a downward arrow on the left labeled Impact falls. Add a note at the bottom reading, Most business cases model only row six and hold the rest constant. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the top row and a muted treatment on the bottom row, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter.
The variables that actually move the number Price per token is the one everybody models and the one that matters least.

Human review time is a real cost and belongs in the same model. An output that needs six minutes of checking has not saved thirty minutes of work; it has saved twenty four and moved the remainder to somebody more expensive. Business cases that count the inference cost and not the review cost are measuring the cheap half.

What local models actually change

The interesting property of running a model on hardware you control is not that it is cheaper per unit. Sometimes it is, and that comparison is unstable enough that anybody quoting you a fixed ratio is quoting a snapshot.

The interesting property is that the cost becomes fixed rather than marginal. Once the hardware exists, an additional call costs electricity and latency rather than money, which changes the shape of what is reasonable to attempt. Under per-call pricing, running a check twice, or running three approaches and comparing, or reprocessing the whole back catalogue because you improved a prompt, are all things you weigh. Under fixed capacity, they are just Tuesday.

That shift matters most for exactly the high volume, low judgment work that dominates real enterprise usage: classification, extraction, routing, summarizing, first-pass reconciliation. Work where the output is checked anyway, where the difficulty is low and the volume is enormous.

There is a second property that is not about money at all. Data that never leaves your control is a different conversation with legal, procurement, and any customer contract with a data residency clause. For some workloads that is the entire argument, and the cost comparison is a footnote.

The honest counterweight: local capacity is capital, it needs operating, and it becomes an asset somebody must maintain and eventually replace. It is a real commitment rather than a free lunch, and organizations that cannot staff it should not pretend the hardware cost is the whole cost.

The routing question

Given both lanes, the design question is which work goes where, and the useful axis is not difficulty. It is tolerance for a wrong answer and volume.

High volume, high tolerance, low judgment work is the natural local lane. Wrong answers are caught downstream, the value is in the aggregate, and the economics reward doing far more of it than a per-call budget would permit.

Low volume, low tolerance, high judgment work is the natural frontier lane. The call count is small, the cost is irrelevant against the stakes, and you want the best available reasoning on the problem.

The dangerous quadrant is high volume, low tolerance, which is where teams reach for the frontier model and then discover the bill. That combination is usually a signal that the process needs redesigning rather than routing: something is being asked to be both cheap and consequential, and it should probably be neither.

What to build regardless of the answer

Build the routing seam on day one, even if it points at a single provider today. An abstraction over model selection costs little at the start and is very expensive to retrofit into a codebase that assumed one vendor's response shape.

Measure per-workload, not per-month. A single invoice tells you nothing actionable. Cost attributed to a workload tells you which use case is quietly consuming the budget, and in my experience it is almost never the one people assume.

Instrument the multiplier. Track calls per completed task. It is the number most likely to drift upward without anybody deciding, and the one most responsive to a day of attention.

And keep the fallback honest. If a provider is unavailable, rate limits you, deprecates the model you built on, or changes its pricing mid-year, what happens? Answering that shapes the architecture more than any current price does, because all four of those are ordinary events rather than tail risks.

The part I will not pretend to know

Anyone offering you a precise crossover point, the volume at which local beats frontier, is offering a number that expires. Model prices have moved repeatedly, open weight capability has moved faster, and hardware moves on its own schedule. A ratio computed today is a fact about today.

What does not expire is the structure: fixed versus marginal cost, the multiplier between requests and calls, tolerance as the routing axis, and the value of being able to change the answer without rewriting the system. Build for that, and the specific numbers become an input you revisit rather than a bet you placed.

Image The switch

Midjourney prompt
painterly editorial illustration, close study of a heavy brass mechanical routing switch mounted on a dark panel, lever set between two marked positions, deep navy metal, warm amber lamplight on the brass, soft teal reflection, tactile and deliberate, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
The switch Routing is worth building early, because it is what lets the decision stay open.
Your specifics would sharpen this (2)

The piece stands on general enterprise truth. Each line below marks a place where a detail only you have would hit harder. Approximations are fine, labelled as approximations.

  • This piece deliberately contains no benchmark numbers, because you have not produced any yet and I will not invent them. When you do have your own measurements, the strongest version of this article is the one that names them.
  • Whether to state plainly that you run local models on your own hardware. It is true and it is differentiating, and I have left it implicit rather than asserting a setup on your behalf.
Art still to generate (3)

Every slot in this piece with no asset yet. Copy a prompt, generate it by hand, commit the file, and its entry disappears from this list.

  1. Two lanes
    Midjourney prompt
    painterly editorial illustration, two parallel channels of light diverging at a switching junction, one narrow and warm, one broad and cool, deep navy surroundings, warm amber in the near channel, soft teal in the far one, brass switching gear and worn rails, quiet engineering order, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw
  2. The variables that actually move the number
    Graph prompt
    Draw a clean editorial ranking diagram titled What actually moves the cost of an enterprise AI workload, ordered by impact, largest at the top. Row one, Call volume and how it grows with adoption. Row two, Retries, retrieval, and multi step reasoning, the hidden multiplier on every call. Row three, Context size per call, driven by how much you send, not how much you need. Row four, Human review time on the output. Row five, Data movement and egress. Row six, Price per token. Draw a downward arrow on the left labeled Impact falls. Add a note at the bottom reading, Most business cases model only row six and hold the rest constant. Style, restrained editorial infographic, deep navy and slate on a warm off white ground, one amber accent on the top row and a muted treatment on the bottom row, thin rules, generous whitespace, sans serif labels, no icons, no gradients, no clutter.
  3. The switch
    Midjourney prompt
    painterly editorial illustration, close study of a heavy brass mechanical routing switch mounted on a dark panel, lever set between two marked positions, deep navy metal, warm amber lamplight on the brass, soft teal reflection, tactile and deliberate, generous negative space, no people, no faces --ar 16:10 --v 7 --style raw