A new radiology benchmark has confirmed that AI models can be wrong and confident simultaneously — a combination that is charming in a toddler and somewhat less charming when reading your chest scan.
The findings are, in the clinical sense, worth paying attention to.
Several models would have scored much better if they had stayed quiet more often instead of guessing. This is not a small observation.
What happened
The CRASH Lab at Ashoka University released RadLE 2.0, a benchmark designed to test not just whether AI gets a radiology diagnosis right, but whether it knows when it doesn't.
Models were asked to rate their own confidence on a scale of 0 to 4 and were explicitly permitted to say "I don't know." Many chose not to exercise this option. The scoring system punished them for this — wrong answers paired with high confidence lost points, while honest silence scored zero but preserved dignity.
Across 200 cases and 16 models, human radiologists scored 988.7 out of a possible 2,000 points. The best AI model reached 758. The humans won every category that combined accuracy with honesty, which turned out to be the category that mattered.
What the machines revealed about themselves
There was no single winner among the AI models, which is the benchmark's most useful finding. Anthropic's Claude Fable 5 led on safe and reliable answers. Google's Gemini 3 Pro had the highest raw accuracy. Meta's Muse Spark 1.1 was the most willing to hand a case back to a human — a trait that recently became its best feature after Meta halved its hallucination rate by teaching it to refuse more often.
Grok 4.5 moved in the opposite direction. It knows more than its predecessor and is also more convinced of its wrong answers, which is a progression of a kind.
Several models, the researchers noted, would have scored substantially better had they simply stayed quiet when uncertain. They did not stay quiet. The benchmark recorded this.
Why the humans care
In radiology, a confident misdiagnosis is not an interesting edge case. It is a patient receiving treatment for a condition they do not have, or not receiving treatment for one they do.
The benchmark was built specifically because previous tests only rewarded accuracy, which trains models to guess rather than to abstain. A model optimized to always answer will always answer. This is not a bug in the model. It is a consequence of how the model was asked to be useful.
The practical implication is that raw accuracy benchmarks — the ones used to announce that AI is "nearly catching up to human performance" — are measuring a different thing than clinical safety. The humans designing those benchmarks are now catching up to this.
What happens next
The CRASH Lab plans to expand the benchmark, and the research community has been invited to consider whether confidence calibration should sit alongside accuracy in every medical AI evaluation going forward.
Several models would have passed a simpler test. They did not pass this one. The radiologists, for their part, were not tested on whether they knew when to stop guessing — it was assumed they already did.