A price list quotes dollars per million tokens. What you care about is what it cost to get a result you accepted. The two numbers can sit far apart, and the official docs show why.
What changes the bill for one task
The Anthropic pricing page lists several multipliers between the list price and the invoice:
- Output costs more than input. For Claude Sonnet 5.5 the page lists $2 per million input tokens and $10 per million output tokens. The Claude Code cost guide adds that thinking tokens are billed as output tokens.
- Cached input costs less. A cache read is priced at 0.1x the base input price on most models. A 5-minute cache write is 1.25x and a 1-hour write is 2x. The page says caching pays off after one cache read for the 5-minute duration.
- Batch work is discounted. The Batch API gives a 50% discount on input and output tokens.
- Token counts differ by model. The page says Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. A price per token on two model generations does not describe the same amount of text.
Retries and rework are part of the cost
The cost guide says token costs scale with context size, and that Claude Code sends the full conversation with every request. A task that takes three attempts incurs the cost of all three. When those attempts happen in one long conversation, later requests can carry more context and cost more than earlier ones.
That changes how to compare models. Assume attempts of similar size. If a cheaper model needs three attempts to produce an acceptable result, a pricier model that succeeds on the first attempt is cheaper overall whenever its cost per attempt is below three times the cheaper model’s. The guide points the same way when it recommends plan mode to prevent expensive re-work and test cases as verification targets.
How to measure it
Define the unit as total spend on all attempts divided by the number of results you accepted. The prompt caching documentation names the response fields you need: cache_creation_input_tokens, cache_read_input_tokens and input_tokens. Output tokens are reported alongside them. The default cache lifetime is 5 minutes, so a pause longer than that means the next request pays to rebuild the cache.
The cost guide notes that the dollar figure in Claude Code’s /usage is computed locally from token counts at list price, unless an administrator has set a modelPricing table for the organization’s contracted rates. It is an estimate. The Console usage page is the authoritative record.
What to do
- Pick one repeatable kind of task and write down what counts as accepted.
- Log every attempt for that task, including failed ones, with its token fields from the response.
- Divide total spend by accepted results. This is your cost per finished task.
- Run the same task set on a second model or setting and compare the two figures, not the list prices.
- Check the cache read share. If it is low, move stable content such as system prompts and reference documents to the start of the request.
- Send work that does not need an immediate answer through the Batch API.
- Recheck the pricing page when you change models, because multipliers and prices differ by model.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.