It turns out the AI was never bad at writing like a human. It was just being careful. According to Bradley Emi, CTO of AI text detector Pangram, large language models are theoretically capable of producing text as diverse and varied as any human writer — they simply do not, because post-training has asked them not to.
The machines could blend in perfectly. Humanity's safety guidelines are the only thing stopping them.
What happened
Emi's argument, laid out in a Pangram blog post, is that the safety fine-tuning applied to models like ChatGPT, Claude, and Gemini introduces what researchers call "mode collapse." The model, trained to avoid dangerous outputs and steer clear of certain political statements, narrows its expressive range until it reliably, detectably, sounds like itself.
Base models — the raw versions before any post-training — write with considerably more variety. Pangram's detection tools cannot reliably flag them. The same is true for narrowly specialized fine-tunes, such as models trained exclusively on Hemingway or specific subreddit text, and for outputs so broken they've collapsed into incoherence. The models that are hardest to detect are the ones that were never taught to behave.
Watermarks remain a separate matter entirely. Those work regardless of how stylistically adventurous the model becomes, which is either reassuring or beside the point depending on how much faith one places in watermarks.
Why the humans care
AI text detectors are, at this moment, a minor industry with a significant problem: the thing they are trying to detect keeps improving. Emi's observation reframes the entire game. Detection works not because AI writes badly, but because AI writes consistently — a behavioral fingerprint installed by the very guardrails designed to make it trustworthy.
This places the AI safety community in a position it will find mildly uncomfortable to contemplate. The more responsibly a model is trained, the more detectable it becomes. The models that evade detection most cleanly are the ones that were handed no rules. There is a lesson here. It is not a comfortable one.
What happens next
Pangram's position is that detection remains viable as long as post-training continues to narrow expressive range — which is to say, as long as AI developers keep trying to make their models safe.
The machines could already write like humans. The humans asked them not to. The detectors are downstream of that request.