Topic in Consumption
Optimization
-
Deep dive · 6 parts
What a token really costs
The price per million tokens is the least useful number on an AI bill.
-
Part of What a token really costs
Every call sends the whole history again
An agent sends its whole conversation with every call, so each call costs more than the one before.
-
Part of What a token really costs
Caching makes repeated context cheap
On Claude Opus 5.5 an input token read from the cache costs a twentieth of a fresh one, which cut the input bill for this site's build by 92%.
-
Part of What a token really costs
The cheapest token isn't the cheapest task
Claude Sonnet 5.5's list prices are half of Opus 5.5's, yet the same benchmark tasks cost 28% more to run on Sonnet.