WorkThe story, in brief

Testing robustness against unforeseen adversaries

OpenAI's new UAR metric exposes a critical blind spot: your AI model's defenses crumble against attacks it's never seen.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

This research addresses a fundamental AI safety and robustness challenge—the gap between tested and real-world adversarial threats. For enterprises deploying classifiers in high-stakes environments, it signals that current robustness claims may be overconfident.

The key facts

8 to know
  1. New metric: UAR (Unforeseen Attack Robustness)

  2. Measures single-model resilience against unanticipated adversarial attacks

  3. Highlights need for evaluation across diverse, unseen attack vectors

  4. Published by OpenAI research team

  5. August 2019 — academic/research contribution to AI safety

  6. Measures single-model performance against unanticipated adversarial attacks

  7. Highlights gap in evaluating robustness across diverse attack scenarios

  8. Dated Aug 22, 2019 — foundational safety research period

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: We’ve developed a method to assess whether a neural network classifier can reliably defend against adversarial attacks not seen during training. Our method yields a new metric, UAR (Unforeseen Attack Robustness), which evaluates the robustness of a single model against an unanticipated attack, and…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work