Procgen Benchmark
OpenAI just released a new RL benchmark. Here's why generalization speed matters more than raw performance.

Why it matters
OpenAI's Procgen Benchmark establishes a new standard for measuring how quickly RL agents learn generalizable skills across procedurally-generated environments—a critical capability for deploying AI systems in unpredictable real-world scenarios.
The key facts
5 to know16 procedurally-generated environments
Focus: generalization and learning speed, not just raw performance
Published by OpenAI in December 2019
Direct application to reinforcement learning agent evaluation
Tests agent capability to transfer learning across novel conditions
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’re releasing Procgen Benchmark, 16 simple-to-use procedurally-generated environments which provide a direct measure of how quickly a reinforcement learning agent learns generalizable skills.