WorkThe story, in brief

AI search agents often confirm what they already know instead of actually researching the web

GPT-5.4 and Kimi K2.6 aren't actually researching. New benchmark reveals AI search agents just confirm what they already know.

Illustration of independent geometric mechanisms passing paper tasks along branching amber tracks.
AI agents and the coordination of work.AI illustration by KeyNews
The KeyNews take

Why it matters

A new academic benchmark (LiveBrowseComp) exposes a fundamental limitation in leading AI search agents: they rely on training data rather than genuine web research, with performance collapsing on recent events. This matters for founders/investors evaluating AI agent reliability for real-world deployment.

The key facts

5 to know
  1. Harbin Institute of Technology published LiveBrowseComp benchmark

  2. Benchmark tests only events from last 90 days

  3. GPT-5.4 and Kimi K2.6 performance significantly degraded on recent-event tasks

  4. Existing model rankings reshuffled when training data advantage removed

  5. AI search agents appear to use web primarily for confirmation rather than discovery

Go to the source

The Decoderthe-decoder.com

Publisher excerpt: Leading AI search agents like GPT-5.4 and Kimi K2.6 don't appear to do much actual research on established benchmarks. They mostly just use the web to confirm what they already learned during training. Researchers at the Harbin Institute of Technology found this using a new time-based benchmark…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work