FrontierFebruary 18, 2026via OpenAI Blog

Introducing EVMbench

Why it matters

OpenAI and Paradigm are establishing a new benchmark for evaluating AI agents on high-stakes security tasks—detecting, patching, and exploiting smart contract vulnerabilities. This signals a major shift in how AI capabilities are measured beyond traditional NLP/reasoning tasks and into specialized domain expertise (blockchain security), suggesting agents are moving from theoretical to production-critical applications.

Key signals

  • EVMbench evaluates AI agents on smart contract vulnerability detection, patching, and exploitation
  • Co-released by OpenAI and Paradigm
  • Focuses on high-severity vulnerabilities in EVM (Ethereum Virtual Machine) contracts
  • Represents expansion of AI benchmarking beyond general reasoning into specialized security domains
  • Indicates movement toward AI agents as tools for specialized professional work (smart contract auditing/security)

The hook

OpenAI just released EVMbench. Here's why smart contract security just became an AI capability race.

OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-severity smart contract vulnerabilities.

The week's key stories, every Friday.

ONE BRIEFING · EVERY FRIDAY · FREE

Free. Unsubscribe anytime.