FrontierThe story, in brief

Evolved Policy Gradients

OpenAI's metalearning breakthrough: agents that generalize to tasks they've never seen before.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Evolved Policy Gradients demonstrates a new approach to training efficiency and task generalization—agents can succeed on novel tasks outside their training distribution, which matters for building more adaptable AI systems.

The key facts

10 to know
  1. OpenAI released Evolved Policy Gradients (EPG), a metalearning approach

  2. EPG evolves the loss function of learning agents

  3. Enables fast training on novel tasks

  4. Agents tested on navigation tasks outside training regime (different object placement)

  5. Demonstrates out-of-distribution generalization capability

  6. Published April 18, 2018

  7. Metalearning approach evolves loss functions of learning agents

  8. Agents generalize to out-of-distribution scenarios (e.g., navigating to object on different side of room)

  9. Published April 18, 2018 — historical research release

  10. OpenAI research contribution

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: We’re releasing an experimental metalearning approach called Evolved Policy Gradients, a method that evolves the loss function of learning agents, which can enable fast training on novel tasks. Agents trained with EPG can succeed at basic tasks at test time that were outside their training regime,…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier