FrontierFebruary 18, 2026via OpenAI Blog
Introducing EVMbench
Why it matters
OpenAI and Paradigm are establishing a new benchmark for evaluating AI agents on high-stakes security tasks—detecting, patching, and exploiting smart contract vulnerabilities. This signals a major shift in how AI capabilities are measured beyond traditional NLP/reasoning tasks and into specialized domain expertise (blockchain security), suggesting agents are moving from theoretical to production-critical applications.
Key signals
- EVMbench evaluates AI agents on smart contract vulnerability detection, patching, and exploitation
- Co-released by OpenAI and Paradigm
- Focuses on high-severity vulnerabilities in EVM (Ethereum Virtual Machine) contracts
- Represents expansion of AI benchmarking beyond general reasoning into specialized security domains
- Indicates movement toward AI agents as tools for specialized professional work (smart contract auditing/security)
The hook
OpenAI just released EVMbench. Here's why smart contract security just became an AI capability race.
OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-severity smart contract vulnerabilities.