How does temperature actually change a model’s output?
At each step, an LLM produces a probability distribution over the next possible token, and temperature reshapes that distribution before a token is sampled. Near zero, the model almost always picks the single highest-probability token, so inference becomes close to deterministic and repeatable across calls. Push it higher and lower-probability tokens get a real chance of being picked, which is what produces more varied phrasing and, past a point, output that drifts off topic or stops making sense. There is no universally correct setting, only the right one for what the output needs to do.
What temperature should a production system actually use?
Near zero for anything that must be consistent or parsed by a program: tool arguments, structured output, classification labels. Higher for creative drafting, brainstorming, or varied phrasing across many generations. Temperature is unrelated to how a model reasons through a problem; it only shapes sampling at the output layer, which is why some reasoning models restrict or ignore it entirely during their internal reasoning steps.