WorkThe story, in brief

Learning from human preferences

OpenAI and DeepMind just solved a core AI safety problem: teaching systems what humans actually want, not what we think we want.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

A foundational safety approach—learning human preferences instead of hand-coded reward functions—addresses a critical gap in AI alignment that remains central to safety debates today. This represents early-stage academic/safety governance work with long-term implications for how AI systems are trained.

The key facts

9 to know
  1. Collaboration between OpenAI and DeepMind safety teams

  2. Algorithm learns human preferences from binary comparisons (A vs B)

  3. Addresses goal specification problem: proxies for complex goals can cause dangerous behavior

  4. Published June 2017 (foundational RLHF-adjacent research)

  5. Focus on safety governance and alignment methodology

  6. Algorithm learns human preferences from pairwise behavior comparisons

  7. Removes need for manually-written goal functions

  8. Published June 2017 — foundational work in RLHF/preference learning

  9. Directly addresses specification gaming and reward misalignment risks

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work