Deep dive · 6 parts
What a token really costs
The price per million tokens is the least useful number on an AI bill.
TL;DR
- Models charge per million tokens of input, everything you send them, and output, everything they generate. Output costs 5 to 8.3 times as much.
- At Claude Opus 5.5’s input and output prices, the session that built this site’s skeleton would have cost $85.76. Its bill was $10.19.
- An agent sends its whole history with every call, so most of its input is repeated. A cache makes the repeats cheap (98.6% of this session’s input came from one), but costs extra when nothing reads it back before it expires.
- 69% of the session’s output was thinking: billed at the output price, mostly never shown.
- A token is a different amount of text in each model, and each model needs a different number of them for the same task. Claude Sonnet 5.5’s tokens cost half as much as Opus 5.5’s, yet a task cost 28% more.
- Compare models by what a task costs, measured on tasks like yours.
cost per task = attempts × Σ (tokens of each type × its price)
I built this site with an AI agent. The session that built its skeleton and its deploy script ran for two hours. It made 92 calls to Claude Opus 5.5, sent it 20.5 million input tokens and got 189K output tokens back.
Multiply those tokens by the prices on the label, $4 per million input tokens and $20 per million output tokens, and you get $85.76. Priced the way the bill is actually worked out, the session came to $10.19.
The gap is what happens between a price list and a bill. Each part below takes one piece of it.
How the numbers were measured
- The session. Claude Code’s transcript records the token usage of every model call. Only those counts are published here, priced at Anthropic’s API list prices.
- Prices come from the providers’ pricing pages, checked on 6 October 2026. Each part links to its sources.
- Token counts of the same paragraph come from each model’s tokenizer, or from the model itself where the tokenizer isn’t public.
- Cost per task comes from Artificial Analysis, which runs the same benchmark on every major model, checked on 6 October 2026.
The parts
- Output costs more than input Every major model charges several times more for a token it generates than for a token you send it.
- Every call sends the whole history again An agent sends its whole conversation with every call, so each call costs more than the one before.
- Caching makes repeated context cheap On Claude Opus 5.5 an input token read from the cache costs a twentieth of a fresh one, which cut the input bill for this site's build by 92%.
- You pay for thinking you never see More than two thirds of the output tokens in this site's build were thinking, paid for at the output price and mostly never shown.
- A token isn't a fixed amount of text The same paragraph is 59 tokens to GPT-5 and 81 to Claude Opus 5.5, so prices per million tokens don't compare directly.
- The cheapest token isn't the cheapest task Claude Sonnet 5.5's list prices are half of Opus 5.5's, yet the same benchmark tasks cost 28% more to run on Sonnet.