Learning from human preferences
A foundational safety approach—learning human preferences instead of hand-coded reward functions—addresses a critical gap in AI alignment that remains central to safety debates today. This represents early-stage academic/safety governance work with long-term implications for how AI systems are trained.
Why it ranks · · Collaboration between OpenAI and DeepMind safety teams · June 2017
Read full story