AgentsAugust 21, 2026via Vercel Blog

How Ora benchmarks every major AI agent on Vercel

Why it matters

A production agent benchmarking platform reveals where agents fail on real websites and how framework choice (Claude Code, ChatGPT, Gemini, eve) affects task completion. Real deployment data for practitioners choosing agent stacks.

Key signals

  • Ora runs agents on live customer websites to measure discovery, navigation, and transaction capability
  • 7% fewer steps to goal; 2x native success rate (tasks completed on customer site vs falling back to web search); 9% more valid endpoints
  • Agents tested: Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, eve (Vercel's framework)
  • 99% of the web estimated as not yet agent-ready by Ora's measure
  • eve (Vercel framework) passed Ora's benchmarks; Ora now uses eve for internal agent infrastructure
  • Sandbox override feature in eve allows Ora to instrument and trace eve agents in their own execution environment
  • 16-person engineering team at Ora ships hundreds of commits per day using coding agents on Vercel stack
  • Prompt-caching fix in eve reduced total cost by ~15% in subsequent benchmarks
  • Platform runs entirely on Vercel: frontend, backend, and agent runtime share deployment, logs, authentication

The hook

Ora benchmarks every major agent framework side-by-side on live websites. The results: 7% fewer steps, 2x native success, 99% of the web still isn't agent-ready.

Ora on Vercel Every harness expects its own infrastructure One platform under every harness Testing eve like any other harness The framework behind Ora's own agents Front end, back end, and agent runtime on one platform Every major agent tested side by side on live sites Hundreds of commits a day

The week's key stories, every Friday.

For practitioners and enthusiasts — free, in your inbox.

Free forever. No spam.