finops.work

What a token really costs · Part 3 of 6

Caching makes repeated context cheap

On Claude Opus 5.5 an input token read from the cache costs a twentieth of a fresh one, which cut the input bill for this site's build by 92%.

An agent sends its whole history with every call, so the model has already seen most of its input. Providers let you cache that part. The first time, the prompt is written to the cache. On later calls, everything up to the first change is read back from it at a fraction of the price.

Cache writes and cache reads are both input, tokens you send. The words describe what the provider does with them: it stores them the first time and reuses them after that.

On most Claude models, a cache read costs a tenth of the input price. On Claude Opus 5.5 it costs a twentieth: $0.20 per million tokens instead of $4. A cache write costs a quarter more than plain input for a cache that lasts five minutes, or twice as much for one that lasts an hour.

In the build session, 98.6% of the input tokens were cache reads. The input cost $6.40. Without caching, the same 20.5 million tokens would have cost $81.97.

What the session's input costOne Claude Code session with Claude Opus 5.5 on 6 October 2026, at Anthropic's list prices.
What the session's input cost Without caching: $81.97; With caching: $6.40. Without caching: $81.97 Without caching $81.97 With caching: $6.40 With caching $6.40

When not to cache

A cache write costs more than plain input, so a cache pays only when it is read back: on Claude, one read pays back a five-minute cache, and two pay back a one-hour cache. It doesn’t pay when:

  • Nothing reads it back in time. A one-off request, or a job that runs once a day, writes a cache that expires unread. Claude’s cache lasts five minutes after it was last used, or an hour at the higher write price. OpenAI’s current models keep it for at least 30 minutes.
  • The start of the prompt changes on every call. A cache matches from the first token up to the first difference. A timestamp at the top of the instructions, or tools listed in a different order, means every call writes and none reads. Keep what changes at the end.
  • Calls come every few minutes, and you pay for an hour. Every read restarts the cache’s clock at no charge, so a five-minute cache that is read every few minutes never expires. A one-hour cache would only cost more.
Go deeper: the arithmetic

The session’s input, priced at Claude Opus 5.5’s list prices, in US dollars:

Input tokens Tokens Per million Cost
Sent fresh 184 $4 <$0.01
Written to the 1-hour cache 295,177 $8 $2.36
Read from the cache 20,196,773 $0.20 $4.04
Total 20,492,134 $6.40
Without caching 20,492,134 $4 $81.97

The break-even: the same prompt sent several times to Claude Opus 5.5, priced as a multiple of sending it once without caching. The first send writes the cache, and every later one reads it.

Times sent No cache Five-minute cache One-hour cache
1 1× 1.25× 2×
2 2× 1.3× 2.05×
3 3× 1.35× 2.1×
10 10× 1.7× 2.45×

Short prompts aren’t cached at all. The minimum is 512 tokens on Claude Opus 5.5 and Sonnet 5.5, 1,024 on OpenAI’s GPT-5.6 and later, and 4,096 on Claude Haiku 4.5, Gemini 3.8 Flash and Gemini 3.1 Pro Preview. Claude processes a shorter prompt without caching and returns no error.

Other providers price caching differently. On OpenAI’s GPT-5.6 and later models, a cache write costs 1.25 times the input price and a cache read a tenth, or a twentieth on GPT-6.1 Sol. Google charges a tenth of the input price for cached tokens on its current Gemini models, and nothing extra for the caching it does on its own. A cache you create yourself also costs storage for every hour it is kept: on Gemini 3.1 Pro Preview, $4.50 per million tokens an hour, so it has to be read more than twice an hour to save money.