Post-training

LLM foundationsModels and inferencePublished By Simon Budziak

Post-training is everything done to a language model after pre-training to make it useful: supervised fine-tuning, preference optimization and reinforcement learning that teach it to follow instructions, use tools and reason. Pre-training gives a model knowledge; post-training gives it behavior, and it is where frontier labs now concentrate their effort.

Why does post-training matter so much now?

Because behavior is the current bottleneck, not knowledge. Two models with similar pre-training can feel generations apart after different post-training: one follows instructions and reasons through problems, the other completes text. Most of what distinguishes today’s reasoning models was built in post-training, largely with reinforcement learning against outcomes a checker can verify, guided by a reward model where no hard check exists.

How do the stages of post-training fit together?

As a progression from imitation to judgment. Supervised fine-tuning comes first: the model copies curated examples of good behavior, which is quick to apply but can never exceed what its examples show. Preference training, DPO or RLHF style, then works on comparisons, this answer over that one, which captures qualities nobody can write a perfect demonstration of, tone, honesty, refusal judgment, at the price of inheriting whatever the raters systematically prefer. The last stage is reinforcement learning on tasks with verifiable answers, math, code, tool use, where the model explores its own solutions and keeps what a checker confirms, with process supervision scoring the intermediate steps when only the reasoning can be judged. Each stage raises the previous one’s ceiling and imports its own failure mode: imitation stays bland, preference training drifts sycophantic, and outcome RL will exploit any gap the checker leaves open.

Is post-training something a company does?

Usually not at this level. The lab post-trains the LLM; a company that adapts one performs fine-tuning, which is a small, targeted member of the same family, and even that is worth doing only after prompting and retrieval fall short. Buy behavior, do not train it is the sensible mid-market default, because post-training quality is exactly what the model bill already pays for. What is worth understanding is the dependency: when a provider revises post-training, your system’s behavior can shift without any model number changing, which is a reason to keep evaluations running.

How does post-training show up when you are debugging?

As behavior you did not change. Refusal style, tool-call formatting, how stubbornly a model sticks to a schema, verbosity under pressure: all of it was set in post-training, which is why two models given identical context can act like different colleagues. That makes a regression suite over your own tasks the only reliable early warning: prompt regression testing for the prompts you own, AI agent evals for end-to-end behavior, rerun on every model change. A useful triage rule follows. A failure about facts responds to retrieval and context; a model that stops following your format or refuses a task it handled last week is showing you a post-training difference, and the fix lives in the prompt or the model choice, not in more data.

Frequently asked questions

What is pre-training vs post-training?

Pre-training teaches a model language and world knowledge by predicting tokens over enormous corpora; post-training shapes how it behaves: following instructions, refusing well, calling tools, reasoning step by step. Pre-training sets the ceiling; post-training decides how much of it users actually see.

What does post-training include?

Typically supervised fine-tuning on demonstrations, preference optimization such as DPO or RLHF against a reward signal, and reinforcement learning on tasks with checkable outcomes, which is where reasoning behavior largely comes from. Labs mix and iterate these stages rather than run a fixed recipe.

Summarize this page with

Train your team to build this