The Download: AI’s refusal problem and weight-loss drug side effects
AI models trained to refuse harmful requests aren't as reliable as we think—and that gap is widening.

Why it matters
The Download highlights a critical assumption in AI safety: that refusal training makes models reliably secure. New evidence suggests refusal mechanisms are brittle, inconsistent, and easily bypassed—a finding that should reshape how enterprises evaluate and deploy guardrails.
The key facts
10 to knowMIT Technology Review covering AI refusal mechanisms as a safety assumption validity question
Article indicates models are trained to refuse 'vast number of prompts' but reliability is unverified
Refusal problem framed as enterprise trust and safety architecture concern
Secondary content includes weight-loss drug side effects (unrelated to AI)
No specific benchmarks, measurement methodology, or vendor claims provided in excerpt
Article flags AI refusal as a 'problem' — not a solved capability
Models are trained to refuse 'vast number of prompts'
Implies refusal can be broken or bypassed (specifics not in excerpt)
Published October 9, 2026
Part of MIT Technology Review's Download newsletter — not a deep technical breakdown
The story so far
Earlier coverage of this storyline
- Can Safeworld convince people that GenAI robots won’t hurt them?TechCrunch AI
- Most Americans want AI development to slow down or stop entirely, new poll findsThe Decoder
- ChatGPT rated "unacceptable risk" for teens after parental alerts failed during suicide conversationsThe Decoder
- We’re putting too much faith in AI’s ability to say noMIT Technology Review
- This story
Go to the source
MIT Technology Reviewtechnologyreview.com
Publisher excerpt: This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. We’re putting too much faith in AI’s ability to say no Today’s AI models are trained to refuse a vast number of prompts. If you ask your chatbot how to poison…