Arena AI Model ELO History
NOBODY TALKING: Everyone benchmarks models in a vacuum. This tracker shows what actually happens to flagship AI performance over weeks—and why your favorite model suddenly feels worse.

Why it matters
A live dashboard tracking historical ELO ratings reveals measurable performance decay in flagship models post-launch, challenging the assumption that model quality remains static. The gap between API benchmarks and consumer UI performance highlights a blind spot in how the industry evaluates real-world model degradation.
The key facts
12 to knowLive ELO tracker visualizes performance lifecycle of flagship AI models over time
Shows measurable performance decay weeks after model launches
Tracks one continuous curve per major AI lab rather than all variants
Identifies gap between API endpoint benchmarks and consumer UI performance (system prompts, safety wrappers, quantized models under load)
Open-source project seeking historical evaluation datasets from consumer web UIs rather than raw APIs
Built to surface both generational jumps and slow performance decays in visualization
Arena AI ELO history tracker visualizes flagship model performance lifecycle
Shows generational jumps and slow performance decay curves per major lab
Identifies gap between API endpoint benchmarks and consumer UI performance
Consumer UIs apply system prompts, safety wrappers, and model quantization under load
Open-source project seeks historical evaluation datasets from consumer web UIs
Published as Hacker News show-hn post (May 2026)
Go to the source
Hacker Newsmayerwin.github.io
Publisher excerpt: Hi HN, I built a live tracker to visualize the lifecycle and performance changes of flagship AI models. We've all experienced the phenomenon where a flagship model feels amazing at launch, but weeks later, it suddenly feels a bit off. I wanted to see if this was just a feeling or a measurable…