Why does an LLM use tokens instead of words?
A model needs a fixed, manageable vocabulary to predict over, and human language has far too many possible words for that to be practical. A tokenizer breaks text into a smaller set of common subword pieces instead: common words often stay whole, rarer words split into a few pieces, and unfamiliar strings like a product code or a typo can fragment into many single characters. That splitting is invisible in a chat interface but directly decides cost, since providers bill per token, not per word.
Where does the token count quietly add up?
Retrieved documents in a RAG system, tool definitions, conversation history, and the model’s own reasoning all consume tokens against the same context window, and a system that reflexively pastes in more text than a question needs pays for every extra token whether it helps the answer or not. Tracking and trimming that spend, rather than assuming a bigger LLM plan absorbs it, is one of the fastest ways to cut a production system’s running cost without touching accuracy.