Reinforcement learning with prediction-based rewards
OpenAI cracked a 30-year-old game AI problem. Here's why it matters for autonomous agents.

Why it matters
OpenAI's Random Network Distillation (RND) demonstrates a breakthrough in curiosity-driven exploration for RL agents, surpassing human performance on a notoriously difficult benchmark. This advances the capability of AI systems to learn in sparse-reward environments without explicit guidance—a critical foundation for real-world agent deployment.
The key facts
10 to knowRandom Network Distillation (RND) method developed
First AI to exceed average human performance on Montezuma's Revenge
Prediction-based reward mechanism for curiosity-driven exploration
Published October 31, 2018 (historical research milestone)
Reinforcement learning capability advancement
Random Network Distillation (RND) method published
First to exceed average human performance on Montezuma's Revenge
Prediction-based reward mechanism for exploration
Published October 31, 2018
Demonstrates RL capability advancement in exploration/curiosity
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’ve developed Random Network Distillation (RND), a prediction-based method for encouraging reinforcement learning agents to explore their environments through curiosity, which for the first time exceeds average human performance on Montezuma’s Revenge.