FalseReject: Reducing overcautiousness in LLMs through reasoning-aware safety evaluation
Amazon just solved the problem nobody talks about: LLMs refusing safe requests. Here's how.

Why it matters
Amazon Science published a novel method to reduce 'overrefusal' in large language models—a critical but underaddressed problem where AI systems reject legitimate requests due to overly conservative safety training. This has direct implications for enterprise LLM deployment and user experience.
The key facts
9 to knowFalseReject: graph-based adversarial method for generating training examples
Targets 'overrefusal' problem in LLM safety evaluation
Reasoning-aware safety evaluation approach
Amazon Science research publication
Addresses balance between safety guardrails and usability in production LLMs
Amazon Science research on LLM safety evaluation
Graph-based, adversarial, agentic method for training data generation
Targets 'overrefusal' problem in language models
Published July 18, 2025
Go to the source
Amazon Scienceamazon.science
Publisher excerpt: Novel graph-based, adversarial, agentic method for generating training examples helps identify — and mitigate — "overrefusal".