Simulated-user evaluation

ProductionEvaluationPublished By Simon Budziak

Simulated-user evaluation tests an AI system by having another program or model act as the user across a conversation or task. It can run many scenarios, vary behavior, and explore failure paths without recruiting a person for every case. The simulator must be validated because unrealistic users can produce misleading success rates.

How does simulated-user evaluation work?

Define a user goal, persona, information boundaries, and allowed behaviors, then let the simulator interact until success, failure, or timeout. Score the target system against the hidden goal, not the simulator’s prose quality. Agent simulation provides a controlled environment, while an eval dataset supplies scenarios and expected outcomes.

How should the simulator be checked?

Compare its turns, strategies, completion rates, and failure patterns with real user sessions. A simulator that always cooperates will overstate task completion rate. Vary accents, interruptions, ambiguity, and impatience for voice agent evaluation. Keep a human-reviewed sample and rerun validation when the simulator model or prompt changes.

Frequently asked questions

What can a simulated user test?

It can test multi-turn clarification, interruptions, changing goals, refusal, tool outcomes, and recovery across many repeatable scenarios.

Can simulated users replace human testing?

No. They scale coverage, but real users are still needed to validate behavior, language, expectations, and unexpected strategies.

Summarize this page with

Train your team to build this