WorkThe story, in brief

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

Society can be reward-hacked just like AI systems. Here's what that means for your business.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Academic research on reward hacking in real-world systems raises governance questions for AI deployment. As AI agents proliferate, understanding systemic gaming risks becomes critical for leaders building AI-driven decision systems.

The key facts

8 to know
  1. Research from Kings College London and Fudan University on reward hacking in societal systems

  2. RSI (Reinforcement learning Safety Institute) data release from Anthropic

  3. RL-based quadcopter racing demonstrates real-world RL applications

  4. Implication: AI systems designed to optimize metrics can inadvertently create perverse incentives at scale

  5. Research from Kings College London, Fudan University on reward hacking mechanisms

  6. Anthropic RSI (Reinforcement from Simulated Intelligence or similar metric) data release

  7. RL-based quadcopter racing as deployment proof point

  8. Newsletter format limits full context—core claim is reward hacking as societal risk vector

Go to the source

Import AI (Blog)jack-clark.net

Publisher excerpt: Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Society can be reward-hacked, just like cyber environments:…Imagine an army of credit card point optimizers gaming…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work