← All writing

Do thinking tokens count as output tokens? Yes, and that is the easy half of the sum

You are turning a model response into a cost, the usage object has more numbers in it than the pricing page has rates, and the whole sum depends on which of them you are allowed to add together.

Yes, thinking tokens count as output tokens. They bill at the output rate and they are already counted inside output_tokens. Adding them on top charges you for the same tokens twice.

The documentation answers this one plainly now

This was worth a page of its own once. It is not any more. Anthropic's page on steering thinking says, under its Pricing heading, that output_tokens "remains the inclusive, authoritative total used for billing" and that output_tokens_details "is a read-only breakdown for observability", with a JSON example immediately above. If you came here to settle an argument, it is settled.

So the rest of this is about the other half of the sum, which the documentation does not save you from, because that mistake is not a misreading. It is an omission.

Reading usage.output_tokens_details.thinking_tokens

The billable figure is usage.output_tokens. The thinking figure sits a level below it at usage.output_tokens_details.thinking_tokens, and the documentation says it is always less than or equal to output_tokens. Subtract it and you have roughly the non-reasoning part of the answer.

Are thinking tokens billed separately? No. There is no separate thinking rate to look up and no separate line to add. The nesting is the only thing that makes it look like its own bucket.

Two practical notes. The documentation does not promise the field is always there, and a turn where the model chose not to think has nothing to report, so read it defensively and default to zero rather than guarding at every call site. And if you are streaming, the documentation says the breakdown appears only on the final message_delta event, which is worth knowing before you conclude the field does not exist.

The published example shows why the double count is not a rounding problem.

Summing all three

input 25 + output 348 + thinking 312 = 685

What that response was billed as

input 25 + output 348 = 373

Those are the documented numbers. Most of that output was thinking, so the mistake nearly doubles the figure. There is no ratio to generalise from one example: the size of the error tracks how hard the model decided to think. The effort setting steers that, but there is no thinking budget to fix it at a number, so the share is not a constant you can hard-code.

The mistake that actually costs money is on the input side

input_tokens is not your input. Once prompt caching is on it counts only the tokens after the last cache breakpoint. The rest of the prompt is reported beside it, in cache_read_input_tokens and cache_creation_input_tokens, and those are additive.

The documented caching example is blunter than anything I could invent. A second request against a warm cache reports a cache read of 3546 tokens and input_tokens of 1062. Take the one field as the input and you have accounted for under a quarter of the prompt the model processed.

The money error is smaller than the token error, because cache reads bill at a fraction of the base input rate. It is not nil. Writes cost more than base input, reads cost less, both have published rates, and the rates differ by model, so this is a lookup rather than something to carry in your head.

Which makes the headline rule longer than the one people repeat. Total billed is input plus output, where input means every input number the usage object reports, cache reads and cache creation included.

The two halves are mirror images, which is why they are worth learning together. Thinking sits inside output and must not be added. Cache reads sit outside input_tokens and must not be dropped. Both come from reading the shape of the JSON as if it were the shape of the bill, and the fix for both is to look up every number in the usage object on the pricing page.

Never collapse input and output into one figure

They bill at different rates, so a single summed tokens number cannot be priced and cannot be taken apart afterwards. Output bills at several times the input rate, so it can carry most of the cost even when input carries most of the tokens. A stored figure that had merged the two would hide that, and with it any chance of saying whether a longer prompt or a longer answer moved the bill.

Keep the thinking count for the same reason, and keep it out of the total for the opposite one. It is the only view you have of whether Tuesday's rise came from longer prompts or from the model working harder on the same ones. Store input, output and thinking as three fields, have the total return input plus output, and write the reason next to it, so a total that looks wrong does not invite somebody to helpfully correct it.

This is the same discipline as pricing a vague ticket: a number you are going to make a decision with has to be built out of the right unit, and no precision afterwards rescues the wrong one.

Record the usage before you parse the response

A response you cannot parse still cost you real money, and those are exactly the calls worth having in the logs, because something went wrong on them. Read the usage before the parse, with the reason in a comment so the ordering survives a tidy-up, and before you pull the text out as well: a reply with no text block at all was still billed.

One structured line per call is enough: operation, model, input, output, thinking. Logs you can query later are all it takes to measure cost per call months afterwards, without a migration.

The text block is not reliably first in the content array

The same documentation is explicit that a turn need not begin with thinking: "A simple factual question may get a direct response with no thinking block at all", and "Don't build application logic that assumes every assistant turn starts with one."

Code that reads content[0].text works right up until the model decides to think. A hard input comes back as thinking then text, an easy one as text alone, so the positional read fails on some responses and not others. The failure is intermittent, and it tracks how hard the model found the input, which makes it miserable to reproduce.

Scan the array for the block whose type is text. That also handles redacted_thinking, a real block type, where positional indexing does not. When there is no text block at all, log the stop_reason and the list of block types you did get, because those two fields turn a mystery into a one-line diagnosis.

Six lines to hold against your own code

  • Total billed is input plus output, where input includes cache reads and cache creation.
  • Thinking is stored, never summed. It is already inside output.
  • Input and output are never collapsed into one figure.
  • Usage is recorded before anything is parsed.
  • One structured log line per call, so cost per call can be measured later without a migration.
  • Content blocks are found by type, never by position.

We had to get this right before we could price anything. The unit we charge for is a clarified issue rather than a model call, and an issue costs one unit however many rounds it takes. That conversion is only ever as good as the per-call number underneath it. What the thing does with the issues themselves is a separate matter.

WhatProblem asks these questions for you

It reads new GitHub issues and asks what is missing, in the issue thread, before anyone on your team has to.

Install from GitHub Marketplace