AgentsThe story, in brief

New Archestra's OpenAPPA Saturates Two Major Security Benchmarks with a 0% Attack Success Rate

0%. Archestra's OpenAPPA stopped every attack in enterprise agent benchmarks—while Claude Code auto mode and Microsoft FIDES leaked data 10% and 31% of the time.

Illustration of a transparent lens revealing connected networks across layers of paper.
Exploring the next frontier of AI research.AI illustration by KeyNews
The KeyNews take

Why it matters

A new open-source security engine claims zero successful prompt-injection and hallucination-driven exfiltration attacks on two production-grade agent benchmarks. The comparative data (vs. Claude Code auto mode and FIDES) is specific and testable, making this relevant to practitioners evaluating agent safety in multi-step workflows.

The key facts

6 to know
  1. Archestra released OpenAPPA, an open-source security engine for agent data exfiltration prevention

  2. Claimed 0% attack success rate on Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench

  3. Claude Code auto mode: 10% attack success rate on same benchmarks

  4. Microsoft FIDES: 31% attack success rate on same benchmarks

  5. Attack vectors: prompt injection and model hallucination-driven exfiltration

  6. Published by InfoQ, no independent third-party validation cited

Go to the source

InfoQ AI/MLinfoq.com

Publisher excerpt: Archestra released OpenAPPA, an open-source security engine designed to stop data exfiltration caused by prompt injection or model hallucination. The team reports zero successful attacks when running security benchmarks Bench-Corp (20 multi-step enterprise workflows) and AgentThreatBench, versus…
Read original report
Back to today's editionMore agents news

Keep reading

Related stories

More from Agents