A new benchmark for evaluating patient-facing health AI agents
PatientAgentBench: the first realistic eval framework for health AI agents actually talking to patients.

Why it matters
As health AI agents move from research to deployment, practitioners need rigorous benchmarks that measure real-world performance — not just model capability. PatientAgentBench fills that gap with synthetic patient interactions, raising the bar for agent reliability in clinical settings.
The key facts
9 to knowPatientAgentBench generates synthetic patient health records and realistic clinical vignettes
Evaluates patient-facing agents through simulated patient-agent conversations
Captures real-world agent behavior (not model-only capability)
Targets deployment readiness for health AI agents
Published by Amazon Science, indicating enterprise focus on agent evaluation
PatientAgentBench generates synthetic patient health records and clinical vignettes for realistic agent evaluation
Includes simulated patient agent in conversation loop—tests actual patient-AI interaction patterns, not isolated capability
Amazon Science release indicates enterprise focus on production-ready agent evaluation for healthcare
Addresses gap in existing benchmarks: most LLM evals don't capture multi-turn patient conversation complexity
Go to the source
Amazon Scienceamazon.science
Publisher excerpt: PatientAgentBench generates a synthetic patient health record, a realistic clinical vignette, and a patient agent that converses with the AI system under evaluation, to capture what a patient-facing agent actually has to do.