NOTE · JULY 2026 · TOKENOMICS

Tokenomics.

> tokens feel free.
> that's exactly why the cost is the blind spot.
> input, output, reasoning — priced differently.
> est. read: 2 min.

◊ THE METER IS ALWAYS RUNNING

You don't pay per call. So you never feel it.

£0.0000 one short session, at list price — and counting

Every token has a price. The interface just never shows you the meter.

01 / THE COSTLESSNESS TRAP

Costless at the point of use is not the same as free.

You type, it answers, nothing is charged in front of you — so the cost feels like zero, and complex work quietly runs up a bill no one is watching. The first rule of tokenomics: the thing that feels free is the thing you stop measuring.

02 / NOT ALL TOKENS COST THE SAME

Input is cheap. Output isn't. And reasoning is the one you can't see.

A token isn't a token. What you send in, what comes back, and what the model thinks to itself on the way are billed at very different rates — and the priciest is invisible.

Input — what you send in baseline ·
Your prompt, documents and context. Processed in a single parallel pass — the cheapest thing on the bill. Big context windows are affordable for a reason.
Output — what comes back typically 3–8× input
Generated one token at a time, each needing its own pass — so it costs several times more than input. The longer and chattier the answer, the steeper the bill.
Reasoning — what it thinks to itself billed as output · hidden
On reasoning models, a long internal chain-of-thought runs before the answer. You never see those tokens — but you pay for them, at the output rate.
> The trap: a visible 500-token answer can quietly consume 3,000+ tokens once reasoning is counted. A reasoning model can bill 5–10× more total tokens than a standard one for the same request — none of it in what you read back.

03 / IT SCALES WITH SUCCESS

The better it works, the more it costs.

Token cost grows with adoption, not with failure. The pilot was cheap because barely anyone used it. Roll it out, watch usage climb — and the bill climbs with it. Success is the expensive case, and it's the one least often modelled.

04 / THE PART THAT ISN'T ABOUT MONEY

When costs spiral, access gets rationed.

There's an equity angle underneath the budget one. When token spend runs away, the reflex is to limit who gets the good models — and a cost problem quietly becomes a fairness problem about who in the organisation is allowed to do their best work with the best tools. Worth naming before it happens, not after.

05 / HOW TO STOP FLYING BLIND

Meter from day one. Model a range, not a point.

Turn the invisible meter back on: instrument usage from the first day, not after the first surprise invoice. Forecast a range rather than a single confident number — spend is still hard to predict, and a point estimate will be wrong. Separate the three token types in your model, because a reasoning-heavy workload prices nothing like a retrieval one. And treat rising cost as a signal of adoption to manage, not a fire to put out. The organisations that stay in control aren't the ones spending least — they're the ones who can see what they're spending while they spend it.

David Kolb advises leaders on AI economics that survive contact with real usage. This note is written to be forwarded — if it's useful to a colleague, send it on.

Rates are indicative and vary by provider and model. Sources: LLM API pricing analyses, 2026 (output ≈ 3–8× input; reasoning billed at output rate).

One thing you can do this week

Compare one month of production spend with what you estimated. Then check what your provider is actually billing you for.

What it won't tell you

Most of the difference isn't the headline token price. It's your implementation — prompt design, context management, retries, caching, orchestration and user behaviour.

If this is live for you.

> Send me the thing you can't get a straight answer on. I'll tell you whether I'm useful.

Email me →