GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
Apple's research shows GRPO works nearly as well in native languages as English—opening reasoning models to billions outside the English-speaking world.

Why it matters
GRPO, the training technique powering reasoning breakthroughs, has been studied almost entirely in English. Apple's large-scale study shows multilingual GRPO closes that gap—a capability shift that changes which populations can benefit from frontier reasoning models.
The key facts
10 to knowLarge-scale empirical study of GRPO across multiple base models and training languages
Native-language reasoning training leaves only small gap to English reasoning performance
Study conducted across multilingual and non-English settings
Research focus: Reinforcement Learning with Verifiable Rewards (RLVR) optimization
Published by Apple Machine Learning Research (credible frontier lab source)
Study covers multilingual and non-English GRPO across multiple base models
Finding: native-language reasoning training leaves only small gap vs. English reasoning
Technique studied: RLVR (Reinforcement Learning with Verifiable Rewards) optimized via GRPO
Scope: wide range of training languages and reasoning-language rewards
Source: Apple Machine Learning Research (peer-reviewed capability research)
Go to the source
Apple Machine Learningmachinelearning.apple.com
Publisher excerpt: Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale…