Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO
90% fewer training steps. Kwai AI's SRPO matches DeepSeek-R1 on math and code—and rewrites the playbook on RL efficiency.

Why it matters
Kwai AI demonstrates a significant breakthrough in reinforcement learning efficiency for LLMs, showing that GRPO-style training can be radically optimized through a two-stage approach with history resampling. This directly challenges the compute requirements and timelines competitors face when training reasoning-capable models.
The key facts
5 to knowSRPO reduces RL post-training steps by 90% vs. baseline GRPO
Matches DeepSeek-R1 performance on math and code benchmarks
Two-stage RL approach with history resampling as core mechanism
Addresses known GRPO limitations in sample efficiency
Published April 23, 2025 on Synced
Go to the source
Synced Reviewsyncedreview.com
Publisher excerpt: Kwai AI's SRPO framework slashes LLM RL post-training steps by 90% while matching DeepSeek-R1 performance in math and code. This two-stage RL approach with history resampling overcomes GRPO limitations. Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO first appeared on Synced.