finops.work

What a token really costs · Part 2 of 6

Every call sends the whole history again

An agent sends its whole conversation with every call, so each call costs more than the one before.

A model remembers nothing between calls. To keep working, an agent sends everything again each time: its instructions, the definitions of its tools, every message so far and every tool result. Then it takes one more step and sends it all again.

The session that built this site’s skeleton made 92 calls. The first one sent 46K tokens: the agent’s instructions and tools, and my request. Every file the agent opened, every command it ran and everything it generated joined the history: the output of one call is input to every call after it. By the last call, each request carried 331K tokens.

Added up, the session sent 20.5 million input tokens: 62 times the length of the conversation it ended with.

Tokens sent with each callOne Claude Code session with Claude Opus 5.5 on 6 October 2026, 92 model calls.
Tokens sent with each call Tokens sent with each of 92 calls, rising from 46K on the first to 331K on the last. 46K 331K Call 1 Call 92
Go deeper: how the history grows

Model APIs are stateless: the provider keeps nothing between calls, so the client sends the conversation again each time. A coding agent such as Claude Code starts each call with its system prompt and the definitions of its tools, then the whole conversation, including every tool result.

In this session the history grew by about 3.1K tokens per call on average, and it never shrank: the agent didn’t compact (summarize) its history. When each call adds roughly the same amount, the total input grows with the square of the number of calls. Here the first 46 calls sent 7.1 million tokens and all 92 sent 20.5 million: doubling the session’s length multiplied its input by 2.9×.

What enters the history stays there, including the model’s own thinking (see part 4). A large file opened early in a session is paid for again, as input, on every call after it, which is why agents open long files in parts and keep what commands print short.

Long sessions can also cross a price line. 33 of this session’s calls sent more than 272,000 tokens. On OpenAI’s GPT-6 models, each of those calls would have cost twice the input price and 1.5 times the output price. Claude charges the same rates up to the end of its context window.

The numbers come from the session’s transcript, which records the usage of every model call; only the counts are published here.