Is the token even the right unit of account?
The honest answer is: the token might not be a unit of account at all. That's a problem, because I've spent eleven posts denominating everything in it.
At CAST26, during the open-season Q&A, someone asked me a simple question. The question was roughly: you keep pricing your decisions in tokens — but what does a token actually cost, and is that number even comparable between models? I said something about output tokens being pricier than input, rate cards, rough ratios. It was accurate and it was useless. The honest answer is: the token might not be a unit of account at all. That's a problem, because I've spent eleven posts denominating everything in it.
What I've been quietly assuming
Nearly every economic claim in this series is priced in tokens. Post 3 argued cost-per-leverage — pay for the tier only where the leverage justifies the tokens. Post 9 said a model you have to ration can't hold the high-frequency seat — rationing measured in tokens against a limit. Post 11 scored the architect seat partly on "cost per call" — call cost being, again, tokens.
Every one of those arguments assumes the token is a stable, comparable unit. The way a euro is. I can say "this costs €4 and that costs €6" and the comparison means something because the euro is the same euro in both hands.
The token isn't the same token in both hands. And once you see that, some of my own confident math gets shakier.
Why the token wobbles as a unit
A unit of account has to do one job: let you compare unlike things on a common scale. The token quietly fails that job in at least three ways.
It isn't standardised across vendors. Every provider tokenises inside its own model context. The same sentence becomes a different number of tokens depending on whose tokeniser is counting, and the base rules only roughly rhyme. "One million tokens" is not a fixed quantity of anything — it's a fixed quantity of that vendor's tokens, which is a different amount of language, and a different amount of thinking, than another vendor's million.
The same count buys different amounts of work. A token spent on a strong reasoning model and a token spent on a cheap worker are priced as if they're on the same scale, but they don't purchase the same thing. One buys a slice of careful judgement; the other buys a slice of fast typing. Counting both in "tokens" flattens a real difference into a fake equivalence.
The count itself is partly the vendor's choice. How verbosely a model thinks, how much hidden reasoning it emits, how it's prompted — all of that moves the token count without moving the outcome you actually wanted. You can pay more tokens for the same result, or fewer, based on decisions that aren't yours.

None of that makes tokens useless. Your invoice is real; the meter runs; within a single model, more tokens genuinely means more spend. What it means is that the token is a billing artifact — a thing the vendor meters to charge you — that I've been treating as a unit of account, a thing you can reason and compare with. Those are not the same, and I conflated them.
What that does to my own argument
Here's the uncomfortable part, and the reason this post is sharp rather than smug. If the unit is wobbly, every cost claim I've made inherits the wobble. "Fable 5 is expensive per call" is a statement about that vendor's token pricing at one moment, not a durable fact. "Opus 5 is cheap enough not to think about" is a comparison across tokenisers that don't line up. I stated those like measurements. They were closer to impressions. So let me separate what survives from what doesn't.
The ranking survives. A worker tier is cheaper than an architect tier; a model behind hard rate limits really can't hold a seat you hit a hundred times a day. Those hold because they don't depend on precise, cross-vendor token math — they're about order of magnitude and about availability, not about the third decimal place.
The precision doesn't. Anywhere I implied a clean numeric comparison — "this call costs X, that one costs Y, therefore" — you should read as directional, not exact. The direction is trustworthy. The decimal isn't. That's a real correction to eleven posts, and I'd rather make it in public than let it sit.
What I'd price in instead
If the token is the wrong denominator, what's the right one? I don't think it's another input unit. Every input unit — tokens, calls, GPU-seconds — has the same disease: it measures what you spent, not what you got. The fix is to move the denominator to the output. Price in units of outcome: cost per shipped feature, per passing test, per review that caught a real defect, per task that made it through the gate without rework.
Outcome units have the property tokens lack — they're comparable across models, because they're defined by your result, not the vendor's meter. "This feature cost me N euros to ship, however the tokens fell out" is a sentence that means the same thing no matter whose tokeniser was running underneath. It also puts the pressure in the right place: not "which model has the cheapest tokens" but "which arrangement of models ships the outcome for the least total cost."

This is also the ground the forthcoming book chapter on "Quality at Machine Speed" stands on. Quality-per-unit-cost only means something if the unit is real. If you're optimising tokens-per-dollar, you can win the metric and lose the outcome — cheaper tokens, more of them, worse result. Outcome-per-cost can't be gamed that way, because the outcome is the thing you actually wanted. The chapter is about designing for that; this post is about noticing the denominator was wrong first.
The honest close
I still don't have the clean answer the CAST26 questioner deserved. If they're reading: you were right, and I've been building on a unit I can't fully defend.
But "the unit might be wrong" turns out to be a better place to argue from than pretending it's settled. It keeps the rankings — which is most of what the practical decisions rest on — and it drops the false precision, which I shouldn't have been trading on anyway. Tokens are transitional scaffolding. Outcomes are the thing. I'll be pricing in outcomes from here.
If you're testing Fable 5 / Opus 5 too, where did it land for you? And the sharper version: have you ever caught yourself comparing token costs across vendors as if the numbers were on the same scale — and what changed when you stopped?
Diagram 1: what "1M tokens" actually buys across three seats — and why the columns don't line up. Diagram 2: input units vs. outcome units — measuring what you spent vs. what you got.
Series: Post 1 — Fable 5 talks to machines better than to people · Show me the receipts · Show, don't tell · Why I took the architect tier out · The RPIQ loop · Where the stack was overkill · BowSmith case study · I lost my favourite tool · Opus Max as architect · Fable 5 came back · CAST26: open season as a quality gate · Opus 5: the middle ground I chose · Post 12 — Is the token even the right unit of account?]