Finding GPT-4’s mistakes with GPT-4
OpenAI trained GPT-4 to catch GPT-4's mistakes. Here's why that matters for model quality.

Why it matters
OpenAI introduced CriticGPT, a GPT-4-based model that identifies errors in ChatGPT outputs during RLHF training. This reveals a critical shift in how frontier labs are approaching model improvement: using AI to scale human feedback collection, a potential bottleneck in training next-gen models.
The key facts
5 to knowCriticGPT is a GPT-4-based model trained to write critiques of ChatGPT responses
Designed to assist human trainers during RLHF (Reinforcement Learning from Human Feedback)
Addresses scalability of human feedback in model training pipelines
Published June 27, 2024 on OpenAI's official blog
Indicates focus on training methodology and feedback optimization rather than raw capability leap
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: CriticGPT, a model based on GPT-4, writes critiques of ChatGPT responses to help human trainers spot mistakes during RLHF