finops.work

What a token really costs · Part 1 of 6

Output costs more than input

Every major model charges several times more for a token it generates than for a token you send it.

Every price list has at least two prices. Input is everything you send to the model: your instructions, the conversation so far, documents, tool results. Output is everything the model generates: its answer, its tool calls and its thinking. Output always costs more. Across the current models from Anthropic, OpenAI and Google, an output token costs 5 to 8.3 times as much as an input token, and for most of them exactly 5 times.

So what a task costs depends on its shape. Classifying a document, or answering a question about a long one, is mostly input. Generating code, reports or long answers is mostly output.

The session that built this site’s skeleton was almost all input: output was 0.9% of its tokens. It was still 37% of the cost.

Output's share of the build sessionOne Claude Code session with Claude Opus 5.5 on 6 October 2026, at Anthropic's list prices.
Output's share of the build session Tokens: 0.9%; Cost: 37%. Tokens: 0.9% Tokens 0.9% Cost: 37% Cost 37%
Go deeper: the prices

List prices in US dollars per million tokens, checked on 6 October 2026:

Model Input Output Output ÷ input
Claude Fable 5.1 $10 $50 5×
Claude Opus 5.5 $4 $20 5×
Claude Sonnet 5.5 $2 $10 5×
Claude Haiku 4.5 $1 $5 5×
GPT-6 Astra $10 $50 5×
GPT-6.1 Sol $2 $10 5×
GPT-6 Luna $0.10 $0.50 5×
Gemini 3.1 Pro Preview $2 $12 6×
Gemini 3.8 Flash $0.75 $3.75 5×
Gemini 3.5 Flash-Lite $0.30 $2.50 8.3×

These are standard prices for prompts below each provider’s long-context threshold. Batch processing halves them. OpenAI charges twice the input price and 1.5 times the output price for prompts over 272,000 tokens, and Google does the same on Gemini 3.1 Pro over 200,000; Anthropic charges the same rates across Claude’s whole context window. Gemini 3.8 Flash’s prices double on 1 January 2027.

Caching adds two more prices, both for input: writing a prompt to the cache costs more than plain input, and reading it back costs much less (see part 3).