Can GRPO be 10x Efficient? Kwai AI’s SRPO Suggests Yes with SRPO
Kwai AI demonstrates a significant breakthrough in reinforcement learning efficiency for LLMs, showing that GRPO-style training can be radically optimized through a two-stage approach with history resampling. This directly challenges the compute requirements and timelines competitors face when training reasoning-capable models.
Why it ranks · · SRPO reduces RL post-training steps by 90% vs. baseline GRPO · Apr 21 – 27, 2025
Read full story