FrontierThe story, in brief

Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO

90% fewer training steps. Kwai AI's SRPO matches DeepSeek-R1 on math and code—and rewrites the playbook on RL efficiency.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Kwai AI demonstrates a significant breakthrough in reinforcement learning efficiency for LLMs, showing that GRPO-style training can be radically optimized through a two-stage approach with history resampling. This directly challenges the compute requirements and timelines competitors face when training reasoning-capable models.

The key facts

5 to know
  1. SRPO reduces RL post-training steps by 90% vs. baseline GRPO

  2. Matches DeepSeek-R1 performance on math and code benchmarks

  3. Two-stage RL approach with history resampling as core mechanism

  4. Addresses known GRPO limitations in sample efficiency

  5. Published April 23, 2025 on Synced

Go to the source

Synced Reviewsyncedreview.com

Publisher excerpt: Kwai AI's SRPO framework slashes LLM RL post-training steps by 90% while matching DeepSeek-R1 performance in math and code. This two-stage RL approach with history resampling overcomes GRPO limitations. Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO first appeared on Synced.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier