Process supervision

LLM foundationsEvaluationPublished By Simon Budziak

Process supervision trains or evaluates a model using feedback on intermediate reasoning steps rather than only the final answer. A supervisor can reward sound steps and flag the point where a solution goes wrong. This creates denser learning signals, but the process labels are expensive and may encode the supervisor's own assumptions.

What does process supervision score?

It may label solution steps, plans, tool choices, calculations, or state transitions. A correct final answer does not excuse an invalid or unsafe path. In reasoning models, step-level feedback can train a reward model to distinguish progress from plausible mistakes.

When is it useful?

Use it when intermediate behavior matters or final outcomes provide too little signal, such as mathematics, tool use, and long agent runs. Prefer observable actions and verifiable checkpoints over speculative labels about hidden reasoning. Trace-based evaluation can inspect recorded steps, while AI agent evals combine trajectory quality with end-to-end task success.

Frequently asked questions

How does process supervision differ from outcome supervision?

Process supervision scores intermediate steps, while outcome supervision scores whether the final result is correct or successful.

Does process supervision require visible chain of thought?

Not necessarily. It can supervise explicit steps, actions, tool traces, or structured intermediate states without exposing hidden reasoning.

Summarize this page with

Train your team to build this