Judge bias is a systematic tendency in an evaluator that changes scores for reasons outside the intended rubric. An LLM judge may favor the first answer, longer wording, familiar style, or outputs from its own model family. Bias can create stable-looking rankings that do not match human or task outcomes.
How can judge bias be detected?
Repeat the same judgment after swapping order, matching length, removing model names, and changing presentation. A score that moves when the underlying quality stays fixed reveals evaluator sensitivity. Research on LLM judges documents position and other judging limitations. Compare results with a human-labeled calibration set.
How should teams mitigate it?
Use explicit rubrics, allow ties, average both orderings in pairwise evaluation, and prefer deterministic checks where ground truth exists. Never treat an LLM as a judge as independent evidence merely because it is a different call. Track agreement and subgroup errors through the AI evaluation harness and revisit the judge when tasks or model versions change.
Frequently asked questions
What are common forms of LLM judge bias?
Common examples include position bias, verbosity bias, style bias, self-preference, and preference for confident wording.
Can a better prompt remove judge bias?
A clear rubric helps, but teams still need perturbation tests and comparison with human labels to measure residual bias.
No advertising or tracking cookies, and our visitor counts are anonymous. The Cal.com booking widget loads only if you allow it. Privacy Policy.
The page itself, anything our host sets to serve and secure it, and the anonymous visitor count. Always on, and none of it stores anything on your device.
The Cal.com booking widget. Left off, a booking link opens the booking page instead of a popup, so you can still book a call.