Testing robustness against unforeseen adversaries
OpenAI's new UAR metric exposes a critical blind spot: your AI model's defenses crumble against attacks it's never seen.

Why it matters
This research addresses a fundamental AI safety and robustness challenge—the gap between tested and real-world adversarial threats. For enterprises deploying classifiers in high-stakes environments, it signals that current robustness claims may be overconfident.
The key facts
8 to knowNew metric: UAR (Unforeseen Attack Robustness)
Measures single-model resilience against unanticipated adversarial attacks
Highlights need for evaluation across diverse, unseen attack vectors
Published by OpenAI research team
August 2019 — academic/research contribution to AI safety
Measures single-model performance against unanticipated adversarial attacks
Highlights gap in evaluating robustness across diverse attack scenarios
Dated Aug 22, 2019 — foundational safety research period
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: We’ve developed a method to assess whether a neural network classifier can reliably defend against adversarial attacks not seen during training. Our method yields a new metric, UAR (Unforeseen Attack Robustness), which evaluates the robustness of a single model against an unanticipated attack, and…