Reward model

LLM foundationsEvaluationPublished By Simon Budziak

A reward model is a learned function that scores model outputs or actions according to labeled preferences, task progress, or desired behavior. Training systems use that score as an optimization signal when direct rules are hard to write. The reward model can also be exploited, so its score is not the same as true quality.

How does a reward model learn?

People or automated systems label outcomes, rank candidate answers through pairwise evaluation, or score intermediate steps. The model learns to predict those preferences at scale. Process supervision supplies step-level signals, while outcome labels focus on the final result.

What can go wrong?

Labels may be inconsistent, narrow, biased, or easy to game. An optimized system will follow the measured reward, including its shortcuts. Test for reward hacking with AI red teaming and outcomes outside the training distribution. Synthetic data can expand coverage but may reinforce the generator’s blind spots. Keep human review and direct task metrics alongside the learned score.

Frequently asked questions

How is a reward model trained?

It commonly learns from ranked output pairs, outcome labels, or step-level judgments that indicate which behavior is preferred.

What is reward hacking?

It occurs when the optimized model finds behavior that earns a high reward score without delivering the intended real-world outcome.

Summarize this page with

Train your team to build this