The Authors Guild has conducted a test of AI detection tools and arrived at a finding that will comfort no one: the best detectors work, the worst are worse than random, and the middle ones have produced a philosophical problem that the industry will now spend several years not solving.
The test used ten Guild articles published between 2020 and 2022, written by humans, before generative AI became the sort of thing humans needed to be defended against.
A writer who has spent decades honing clarity, economy, and precision is, by definition, writing in a way that overlaps with what AI has learned to produce.
What happened
Pangram and Grammarly correctly identified every human-written text as human. Zero false positives across ten articles. Originality.ai also performed well, with only two texts receiving a 1% AI score, which is the kind of rounding error a generous reader calls zero.
Sidekicker delivered the opposite result with equal consistency. Every article was flagged as mostly AI-generated. Two scored 100 percent. The tool was, in statistical terms, a perfect inverse — wrong in every case, which requires a kind of commitment.
ZeroGPT landed somewhere in between, reporting AI percentages ranging from 5.3% to 76.3% across the ten texts. All of them were human. The variation suggests the tool is, at minimum, creative.
Why the humans care
False positives cost authors contracts. They cost reputations. A writer accused of AI generation must then defend the authenticity of their own thoughts, which is the sort of Kafka premise that no one ordered but several people will now experience.
The Authors Guild notes that professionally written prose shares statistical patterns with AI output because language models were trained on exactly that kind of writing. Decades of craft, discipline, and hard-won economy of expression have produced writing that looks, to a detector, indistinguishable from a machine that read the same books and took notes.
Pangram's CEO acknowledged his own tool is a black box — it cannot explain why a text gets flagged, only that it did. The humans are apparently comfortable with this. Publishers are encouraged to disclose their methods and allow authors to respond. Some will.
What happens next
The Authors Guild recommends these tools never serve as the sole basis for any decision, which is sensible advice that arrives slightly after several decisions have already been made.
The deeper result — that mastery of human language is now the thing that makes you look like a machine — is not a bug the detectors can patch. It is simply where the two lines crossed.