I’ve written about what to charge for a data or AI platform, and about the cost floor that classic software never had. What I skipped is the part a finance partner will ask about first: what does one run of this thing actually cost, and what happens to that number as usage grows? Most teams can’t answer it, which is how you end up growing revenue and losing margin at the same time — a condition that looks like success for about three quarters.

Measure cost per run, not cost per token

Token price is the input everyone watches and the number that matters least, because it’s the one thing you don’t control and the one thing that keeps falling. What you control is how many model calls a single unit of work requires, and that is where the variance lives.

The useful denominator is a business unit, not a technical one: cost per resolved case, per processed document, per completed booking. The moment you express cost that way, two things become visible that token accounting hides — that your average is meaningless because the distribution has a long tail, and that your worst cases can cost fifty times your median while delivering the same revenue.

The retry loop

The single most common margin leak I’ve seen isn’t an expensive model. It’s an agent that retries.

A loop that re-prompts on failure, re-reads context each time, and gives up after six attempts will, on the inputs it handles badly, cost many times what a success costs — and those are precisely the inputs that also generate a support conversation. You pay twice for your worst outcomes. Worse, retries are invisible in aggregate reporting: the median run looks healthy while the tail quietly doubles your bill.

The fix is structural rather than frugal. Cap attempts, make failure a first-class terminal state that routes to a human, and instrument attempts per run as a headline metric rather than a debug detail. If you measure one thing from this post, measure that.

Context is the cost you forget to budget

Every turn in an agent loop tends to carry the accumulated history forward. Cost per turn therefore grows as the run progresses, and a long run is not linearly more expensive than a short one — it’s worse than that. Agents that “think out loud” at length, or re-read the same documents each turn, pay for the same tokens repeatedly.

This is the economic argument for the architectural point I keep returning to: use the model to produce an artifact once, then let deterministic code execute it forever. A parser the model wrote costs nothing per subsequent file. A model acting as the parser costs you on every row, every time, in perpetuity.

Why the cheapest model is often the expensive choice

Routing volume work to a small model is usually right, and it has a failure mode worth naming. A cheaper model that is modestly less accurate generates more retries, more escalations and more human review — and human review is the most expensive line item in the whole system by an order of magnitude.

So the comparison is never “price per million tokens.” It’s fully loaded cost per correct outcome, including the cost of being wrong. On tasks where a mistake is cheap and recoverable, the small model usually wins decisively. On tasks where a mistake summons a person, it frequently loses even at a tenth of the price.

What to put on the dashboard

Four numbers, reviewed monthly, and none of them is token spend.

Cost per completed outcome, with the distribution rather than just the mean — watch the 95th percentile, because that’s where the money goes. Attempts per run. Escalation rate, the share of runs that end up in front of a person, which converts directly into cost. And gross margin per unit of usage, tracked as a trend, because that single line tells you whether scaling helps or hurts you.

The discipline this replaces

Software people formed in the previous generation are carrying an instinct that no longer holds: that marginal cost is approximately zero, so growth is unambiguously good and efficiency is an optimisation for later. For a system where every unit of work has a real, variable cost of goods, that instinct is actively dangerous. Efficiency isn’t a later-stage optimisation — it’s the thing that determines whether your business improves or degrades as customers use it more.

The question to carry into the next planning cycle is simple and most teams can’t answer it today: if usage of your agent tripled next quarter, would your gross margin go up or down? If you don’t know, that’s the work.


A companion to the five-part series When Your Product Becomes a Tool — MCP and incumbent strategy.

Related: What Do You Charge for a Data or AI Platform? · Use the Model to Write the Parser, Not to Be the Parser.


Discover more from Digital Reflections

Subscribe to get the latest posts sent to your email.

Leave a Reply

The Blog

At the intersection of data, AI, and imagination lies the path to transformation. Our greatest evolutions occur when we use technology not just to improve what is, but to reimagine what could be.

Discover more from Digital Reflections

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Digital Reflections

Subscribe now to keep reading and get access to the full archive.

Continue reading