LLMs could write like humans but post-training guardrails make their text detectable
Base models write like humans. RLHF kills the variety. Here's what post-training safety constraints actually cost.

Why it matters
A technical claim about how safety training narrows LLM expressiveness — relevant to practitioners understanding model behavior and to enthusiasts tracking the frontier labs' training trade-offs.
The key facts
9 to knowPost-training guardrails narrow expressive range in LLMs
Base models (pre-RLHF) exhibit far more stylistic variety
Safety constraints make LLM text detectably machine-generated
Source: Pangram CTO Bradley Emi
Implication: detectability and stylistic uniformity are byproducts of safety training, not capability limits
Bradley Emi (Pangram CTO) claims base models show far greater stylistic variety than post-trained versions
Post-training guardrails narrow expressive range detectably
Safety constraints may be limiting capability, not revealing inherent model limitations
Implication: LLM-generated text detectability may be a feature of alignment, not a fundamental model behavior
Go to the source
The Decoderthe-decoder.com
Publisher excerpt: LLMs don't write in a recognizable style because they can't do better. Post-training and safety guardrails sharply narrow their expressive range, argues Pangram CTO Bradley Emi. Base models without these constraints already write with far more variety.