Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study finds
208,000 participants. One finding: making AI helpful breaks its ability to act human. And it's getting worse each generation.

Why it matters
A large-scale empirical study reveals a fundamental trade-off in LLM design: helpfulness training degrades human behavior simulation, with implications for AI reliability in social modeling, research, and deployment decisions.
The key facts
5 to knowStudy scale: 208,000 participants, 26 million responses
Finding: RLHF/helpfulness training reduces human behavior replication accuracy
Degradation compounds across model generations
Persona injection (demographic profiles) provides minimal improvement for individual predictions
Suggests fundamental architectural tension between alignment and behavioral fidelity
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: A large-scale study covering 208,000 participants and 26 million responses shows that the very training that turns language models into helpful chatbots weakens their ability to replicate human behavior. The effect gets worse with each new model generation. Even the popular persona trick, feeding…
