Solving math word problems
OpenAI's math system beats GPT-3 by 2x. Here's why that matters for reasoning.

Why it matters
OpenAI demonstrates significant capability advancement in mathematical reasoning—a key benchmark for LLM progress. The comparison to child-level performance signals movement toward more robust reasoning systems, though the 55% accuracy highlights remaining gaps in real-world problem-solving.
The key facts
9 to knowOpenAI system achieves ~90% of 9-12 year old performance (55% vs 60%)
Nearly 2x accuracy improvement over fine-tuned GPT-3
Grade school math word problems as benchmark
Published October 29, 2021 (historical research announcement)
Dataset sourced from real child testing
New system achieves ~90% of child performance (55% vs 60% baseline)
2x accuracy improvement over fine-tuned GPT-3
Tested on grade school math word problems
Published October 2021 (archived technical milestone)
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’ve trained a system that solves grade school math problems with nearly twice the accuracy of a fine-tuned GPT-3 model. It solves about 90% as many problems as real kids: a small sample of 9-12 year olds scored 60% on a test from our dataset, while our system scored 55% on those same problems.