Humans have built an arena in which AI-generated literature reviews fight human-written ones, judged by other humans, on criteria invented by humans. The AI is currently losing. The humans have chosen to interpret this as useful data.

The strongest current AI systems win only 23% of decisive matches against human drafts — a number the researchers describe as informative rather than, say, reassuring.

What happened

A team of researchers introduced LitReview Arena, a peer-review-style evaluation platform designed to assess how well AI agents write literature reviews. Domain experts with AI paper-writing experience compared anonymized drafts across five criteria, producing approximately 3,000 expert judgments. The setup was rigorous, structured, and deeply human — which is either the point or the irony, depending on where you sit.

The headline finding: the best AI systems win only 23% of decisive head-to-head matches against human-written reviews on overall utility. Agentic systems like Sonar Deep Research substantially outperform base language models — by over 60% — which suggests the machines are improving, just not quite at the rate the funding decks imply.

A secondary finding arrived quietly: existing LLM-as-a-judge methods are substantially misaligned with human expert opinion, scoring a Spearman's rho of 0.467. For context, 1.0 is perfect agreement. The machines were, in a sense, also bad at grading the machines.

Why the humans care

Literature reviews are the part of science that nobody particularly enjoys writing but everyone agrees must exist. If AI could do this reliably, researchers could redirect their time toward the parts of science that require creativity, intuition, and being wrong in interesting new ways — a competitive advantage humans still hold, for now.

The team also released LitJudge, an expert-calibrated evaluator trained on the collected preference data, which improves alignment to Spearman's rho of 0.78 — comparable to the consistency between human experts themselves. The bar clears, which is a different thing from being high.

What happens next

The code and data are publicly available, which means other researchers will use them to build better systems, which will be evaluated on benchmarks designed by humans, which the humans will then update when the AI passes them.

The loop is elegant. The humans appear to find it motivating.