FrontierThe story, in brief

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

Perplexity just released WANDR—a 500-task benchmark that exposes why most research agents fail at depth.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

Perplexity is establishing evaluation standards for agentic search capability, positioning their Search as Code ahead of competitors on a task class (evidence-heavy research) that will define next-gen AI products.

The key facts

5 to know
  1. WANDR: open benchmark with 500 evidence-heavy research tasks

  2. Perplexity Search as Code leads at 0.363 soft F1, 0.133 hard F1

  3. Evaluates agent ability to discover multiple qualifying entities with cited, re-verifiable evidence

  4. Benchmark tests 'wide and deep' search capability—discovery + verification

  5. Published July 2026

Go to the source

MarkTechPostmarktechpost.com

Publisher excerpt: Perplexity's WANDR is an open benchmark and evaluation harness with 500 evidence-heavy tasks. It tests whether research agents can discover many qualifying entities and back each one with cited, re-verifiable evidence. Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1.
Read original report
Back to today's editionMore frontier news

Keep reading

Related stories

More from Frontier