AI search agents often confirm what they already know instead of actually researching the web
GPT-5.4 and Kimi K2.6 aren't actually researching. New benchmark reveals AI search agents just confirm what they already know.

Why it matters
A new academic benchmark (LiveBrowseComp) exposes a fundamental limitation in leading AI search agents: they rely on training data rather than genuine web research, with performance collapsing on recent events. This matters for founders/investors evaluating AI agent reliability for real-world deployment.
The key facts
5 to knowHarbin Institute of Technology published LiveBrowseComp benchmark
Benchmark tests only events from last 90 days
GPT-5.4 and Kimi K2.6 performance significantly degraded on recent-event tasks
Existing model rankings reshuffled when training data advantage removed
AI search agents appear to use web primarily for confirmation rather than discovery
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: Leading AI search agents like GPT-5.4 and Kimi K2.6 don't appear to do much actual research on established benchmarks. They mostly just use the web to confirm what they already learned during training. Researchers at the Harbin Institute of Technology found this using a new time-based benchmark…