Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Perplexity just released WANDR—a 500-task benchmark that exposes why most research agents fail at depth.

Why it matters
Perplexity is establishing evaluation standards for agentic search capability, positioning their Search as Code ahead of competitors on a task class (evidence-heavy research) that will define next-gen AI products.
The key facts
5 to knowWANDR: open benchmark with 500 evidence-heavy research tasks
Perplexity Search as Code leads at 0.363 soft F1, 0.133 hard F1
Evaluates agent ability to discover multiple qualifying entities with cited, re-verifiable evidence
Benchmark tests 'wide and deep' search capability—discovery + verification
Published July 2026
Go to the source
MarkTechPostmarktechpost.com
Publisher excerpt: Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1.