WorkThe story, in brief

The Download: AI’s refusal problem and weight-loss drug side effects

AI models trained to refuse harmful requests aren't as reliable as we think—and that gap is widening.

Illustration of two anonymous hands arranging task cards around an amber tool on a shared desk.
People, judgement and the changing nature of work.AI illustration by KeyNews
The KeyNews take

Why it matters

The Download highlights a critical assumption in AI safety: that refusal training makes models reliably secure. New evidence suggests refusal mechanisms are brittle, inconsistent, and easily bypassed—a finding that should reshape how enterprises evaluate and deploy guardrails.

The key facts

10 to know
  1. MIT Technology Review covering AI refusal mechanisms as a safety assumption validity question

  2. Article indicates models are trained to refuse 'vast number of prompts' but reliability is unverified

  3. Refusal problem framed as enterprise trust and safety architecture concern

  4. Secondary content includes weight-loss drug side effects (unrelated to AI)

  5. No specific benchmarks, measurement methodology, or vendor claims provided in excerpt

  6. Article flags AI refusal as a 'problem' — not a solved capability

  7. Models are trained to refuse 'vast number of prompts'

  8. Implies refusal can be broken or bypassed (specifics not in excerpt)

  9. Published October 9, 2026

  10. Part of MIT Technology Review's Download newsletter — not a deep technical breakdown

The story so far

Earlier coverage of this storyline

  1. Can Safeworld convince people that GenAI robots won’t hurt them?TechCrunch AI
  2. Most Americans want AI development to slow down or stop entirely, new poll findsThe Decoder
  3. ChatGPT rated "unacceptable risk" for teens after parental alerts failed during suicide conversationsThe Decoder
  4. We’re putting too much faith in AI’s ability to say noMIT Technology Review
  5. This story

Go to the source

MIT Technology Reviewtechnologyreview.com

Publisher excerpt: This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. We’re putting too much faith in AI’s ability to say no Today’s AI models are trained to refuse a vast number of prompts. If you ask your chatbot how to poison…
Read original report
Back to today's editionMore work news

Keep reading

Related stories

More from Work