Why does inference cost money on every single call?
Training happens once, expensively, and produces a fixed set of weights. Inference happens every time anyone uses the model, and each call recomputes predictions across billions of parameters for every token it reads and writes. That per-call cost is why a chatty AI agent looping through many steps burns real spend on cycles that move nothing forward, unlike training’s one-time bill.
What actually determines how fast or expensive inference is?
Mostly the size of the input and output measured in tokens, plus how long the user waits before the first token arrives, the responsiveness measure called time to first token: filling the context window with retrieved documents, tool definitions, and history all adds to the bill. Streaming inference exposes output earlier without necessarily reducing total computation. The sampling settings matter too, though differently; temperature changes how the model samples its next token but not how much compute that token costs. Production systems that watch this closely, trimming context, caching repeated prefixes through prompt caching, or serving quantized weights on cheaper hardware, are the ones that keep inference costs from quietly outrunning the value the system delivers, the same discipline behind halving Claude Code’s token usage.