OpenAI and Anthropic share findings from a joint safety evaluation
OpenAI and Anthropic just evaluated each other's models. Here's what they found about alignment, jailbreaking, and hallucinations.

Why it matters
Cross-lab safety collaboration between two AI leaders sets a governance precedent and reveals real-world model vulnerabilities—critical for boards weighing AI deployment risk and regulatory compliance.
The key facts
5 to knowFirst joint safety evaluation between OpenAI and Anthropic
Testing domains: misalignment, instruction following, hallucinations, jailbreaking
Cross-lab collaboration model for AI safety benchmarking
Published findings available publicly (governance transparency)
Benchmarking approach includes vulnerability assessment across both organizations' models
Go to the source
OpenAI Blogopenai.com
Publisher excerpt: OpenAI and Anthropic share findings from a first-of-its-kind joint safety evaluation, testing each other’s models for misalignment, instruction following, hallucinations, jailbreaking, and more—highlighting progress, challenges, and the value of cross-lab collaboration.
