Learning to summarize with human feedback
OpenAI just cracked the code on training models with human feedback. Here's why that matters for everything that comes next.

Why it matters
This foundational research on RLHF (reinforcement learning from human feedback) became the training methodology behind ChatGPT and modern LLMs. It's a capability breakthrough that shaped how the entire industry trains alignment into models.
The key facts
10 to knowRLHF training methodology applied to summarization tasks
Published September 2020 — predates ChatGPT by 2+ years
Foundational research that enabled ChatGPT's alignment approach
Demonstrates human feedback can improve model quality at scale
Language model training approach innovation
RLHF applied to language model summarization task
Demonstrates human feedback as training signal for alignment
Published September 2020 (pre-ChatGPT era, foundational research)
Technique later scaled to become core of InstructGPT and GPT-3.5
Established playbook adopted across Claude, Gemini, and other frontier models
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’ve applied reinforcement learning from human feedback to train language models that are better at summarization.