FrontierThe story, in brief

Arena AI Model ELO History

NOBODY TALKING: Everyone benchmarks models in a vacuum. This tracker shows what actually happens to flagship AI performance over weeks—and why your favorite model suddenly feels worse.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A live dashboard tracking historical ELO ratings reveals measurable performance decay in flagship models post-launch, challenging the assumption that model quality remains static. The gap between API benchmarks and consumer UI performance highlights a blind spot in how the industry evaluates real-world model degradation.

The key facts

12 to know
  1. Live ELO tracker visualizes performance lifecycle of flagship AI models over time

  2. Shows measurable performance decay weeks after model launches

  3. Tracks one continuous curve per major AI lab rather than all variants

  4. Identifies gap between API endpoint benchmarks and consumer UI performance (system prompts, safety wrappers, quantized models under load)

  5. Open-source project seeking historical evaluation datasets from consumer web UIs rather than raw APIs

  6. Built to surface both generational jumps and slow performance decays in visualization

  7. Arena AI ELO history tracker visualizes flagship model performance lifecycle

  8. Shows generational jumps and slow performance decay curves per major lab

  9. Identifies gap between API endpoint benchmarks and consumer UI performance

  10. Consumer UIs apply system prompts, safety wrappers, and model quantization under load

  11. Open-source project seeks historical evaluation datasets from consumer web UIs

  12. Published as Hacker News show-hn post (May 2026)

Go to the source

Hacker Newsmayerwin.github.io

Publisher excerpt: Hi HN, I built a live tracker to visualize the lifecycle and performance changes of flagship AI models. We've all experienced the phenomenon where a flagship model feels amazing at launch, but weeks later, it suddenly feels a bit off. I wanted to see if this was just a feeling or a measurable…
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier