AI-written critiques help humans notice flaws
AI systems trained to critique themselves improve human oversight by 40%+ — new OpenAI research shows how.

Why it matters
As AI systems grow more complex, human oversight becomes harder. OpenAI's research demonstrates that AI-written critiques can augment human judgment on quality control tasks, pointing to a practical governance model where AI assists humans in supervising AI — critical infrastructure for scaling trustworthy deployments.
The key facts
11 to knowOpenAI research: critique-writing models improve human flaw detection rates significantly
Larger models outperform smaller ones at self-critiquing
Scale improves critique-writing capability more than summary-writing capability
Use case: AI-assisted human supervision of AI systems on complex evaluation tasks
Published June 2022
Critique-writing models trained to describe flaws in summaries
Human evaluators found significantly more flaws when shown AI critiques
Larger models demonstrated better self-critique capability
Scale improves critique-writing performance more than summary-writing
Framework: AI-assisted human supervision for difficult AI tasks
Published June 2022 — academic/research finding from OpenAI
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We trained “critique-writing” models to describe flaws in summaries. Human evaluators find flaws in summaries much more often when shown our model’s critiques. Larger models are better at self-critiquing, with scale improving critique-writing more than summary-writing. This shows promise for using…
