WorkThe story, in brief

OpenAI and Anthropic share findings from a joint safety evaluation

OpenAI and Anthropic just evaluated each other's models. Here's what they found about alignment, jailbreaking, and hallucinations.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

Cross-lab safety collaboration between two AI leaders sets a governance precedent and reveals real-world model vulnerabilities—critical for boards weighing AI deployment risk and regulatory compliance.

The key facts

5 to know
  1. First joint safety evaluation between OpenAI and Anthropic

  2. Testing domains: misalignment, instruction following, hallucinations, jailbreaking

  3. Cross-lab collaboration model for AI safety benchmarking

  4. Published findings available publicly (governance transparency)

  5. Benchmarking approach includes vulnerability assessment across both organizations' models

Go to the source

OpenAI Blogopenai.com

Publisher excerpt: OpenAI and Anthropic share findings from a first-of-its-kind joint safety evaluation, testing each other’s models for misalignment, instruction following, hallucinations, jailbreaking, and more—highlighting progress, challenges, and the value of cross-lab collaboration.
Read original report
Back to today's editionMore work news

The wider picture

View all
Illustration of a transparent lens revealing connected networks across layers of paper.
AI illustration by KeyNews
Work01

The next AI advantage isn’t a bigger model — it’s a better semantic layer

As agentic AI moves from pilots to production, semantic alignment—shared, machine-readable business definitions—has become a prerequisite for safe deployment. This is a people and process problem, not a model problem, and it requires organizational work that most data teams aren't yet structured to do.

CIO
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work02

I Built AI Clones of My Coworkers. Things Got Weird

As AI personalization deepens in workplace tools, the line between mimicry and utility blurs. This is a cautionary tale about how AI-as-coworker products could amplify personality performance over actual work.

Wired AI
Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
AI illustration by KeyNews
Work03

‘Rising threats and under-resourcing for cybersecurity is taking a toll on the people tasked with managing it’: Cyber teams are being pushed to breaking point – and AI is doing little to alleviate strain

AI is being marketed as a cybersecurity force multiplier, but practitioners report it's not reducing workload or burnout in under-resourced teams. This is a workforce reality check on AI's promised impact on a critical profession.

ITPro