finops.work

What a token really costs · Part 4 of 6

You pay for thinking you never see

More than two thirds of the output tokens in this site's build were thinking, paid for at the output price and mostly never shown.

Reasoning models think before they answer: they generate intermediate steps first, then the answer. Thinking is generated one token at a time like any other output, and it is billed as output.

In the session that built this site’s skeleton, the model generated 189K output tokens. 130K of them, 69%, were thinking. The rest were the replies I saw and the tool calls that edited files and ran commands.

At the output price, the thinking cost $2.60: 26% of what the whole session cost. And that was only the first time it was paid for. Claude Opus 5.5 keeps earlier thinking in the conversation, so it travels with the history and is billed again, as input, on every call that follows.

Output tokens in the build sessionOne Claude Code session with Claude Opus 5.5 on 6 October 2026, 92 model calls.
Output tokens in the build session Thinking: 130K; Replies and tool calls: 59K. Thinking: 130K Thinking 130K Replies and tool calls: 59K Replies and tool calls 59K
Go deeper: how thinking is billed

Anthropic’s documentation is explicit: the tokens Claude spends reasoning “are billed as output tokens, even when the thinking text isn’t returned to you”. On Claude Opus 5.5 the API returns no thinking text by default. You can ask for a summary, but you are “charged for the full thinking tokens generated by the original request, not the summary tokens”.

Claude Opus 4.5 and later models keep the thinking from earlier turns in context. Like the rest of the conversation history, it is billed as input tokens on every later call; in an agent’s loop, mostly at the cache-read price.

The other providers bill the same way. OpenAI bills reasoning tokens as output tokens, and Google’s price list gives each Gemini model an “output price (including thinking tokens)”.

The counts come from the session’s transcript. Each model call records its output tokens and, separately, how many of them were thinking (output_tokens_details.thinking_tokens).