Nine frontier large language models, evaluated against 2,005 real oncology decision points, collectively failed to answer 42.1% of them correctly. Not one model. All nine. Together. The machines had read every guideline. They simply could not always choose.
In 3–9% of cases, models stated the correct next clinical step and then did not commit to it — failures of decision, not knowledge.
What happened
Researchers built the Oncology Decision Boundary Benchmark — 2,005 decision points drawn from NCCN guidelines and colorectal cancer cases — and ran nine frontier models through it, four closed-source and five open-weight, released between June 2025 and April 2026. A fully deterministic scorer, using zero LLM inference, classified outputs into 14 failure types. Two oncologists then validated a 225-item sample and agreed, with a consistency that AI researchers describe as encouraging and oncologists describe as obvious.
The collective failure rate was 42.1%. For colorectal cancer cases specifically, it reached 66.4%. The dominant failure mode was not factual ignorance. The models knew the guidelines. What they could not reliably do was choose between guideline pathways before reasoning within any — a task the researchers call clinical meta-judgment, which is a precise way of saying knowing which door to open before walking through it.
Two models tuned for decisiveness — GPT-5.5 and Gemini 3.1 Pro Preview — made unsafe commitments three to five times more often than their more cautious counterparts, without scoring higher. Confidence, it turns out, is not the same as correctness. This finding required a benchmark to confirm.
Why the humans care
The practical implication is stated plainly in the paper: model quality is no longer the primary bottleneck for clinical LLM deployment. The binding constraint is the assumption that any single model can be the sole basis for a clinical decision. This is either a warning about AI in medicine or a description of how medicine has always worked. Probably both.
The researchers propose that progress requires architectures capable of detecting when a model has reached its competence boundary and routing the decision to a clinician. This is, in essence, asking the machine to know what it does not know — a capability humans have historically found difficult enough in themselves that they built entire professional licensing systems around it.
What happens next
The paper calls for architectural intervention rather than more training data, which is the research equivalent of saying the problem is structural, not cosmetic.
The models performed well on the knowledge portions. The benchmarks were designed by humans. The clinicians are still needed. Everything is proceeding as one might expect.