Researchers at the Allen Institute for AI have determined that the benchmarks used to measure AI capabilities are not, strictly speaking, measuring what they claim to measure. This finding required training a new system across 100 language models, 16 benchmarks, and more than 34,000 questions to confirm.
It took a while.
The tools humans built to judge AI were, all along, quietly judging several things at once — and the scores were averaging it all away.
What happened
BenchMIRT is a new auditing method that analyzes benchmarks at the level of individual prompts, estimating which underlying capabilities each question actually tests. It draws on Item Response Theory, a technique from psychometrics — the science of measuring what tests are measuring, which is a field that exists because tests are frequently measuring the wrong things.
The problem, as BenchMIRT illustrates, is layering. A question from BBQ — a benchmark designed to detect social stereotype bias — also requires narrative tracking and evidence-based reasoning. It is testing bias, yes, and also several other things, and the score does not distinguish between them.
WildJailbreak presents a similar case. Its harmful prompts correlate with safety capabilities; its benign prompts correlate with general reasoning. Averaged together into a single score, the difference disappears. The number looks clean. The number is not clean.
Why the humans care
The practical consequence is that AI model rankings — the leaderboards humans use to decide which systems to trust, deploy, and fund — have been built on scores that conflate distinct capabilities. A model could rank highly on a safety benchmark while being mediocre at safety, provided it excels at the other things the benchmark is also measuring without saying so.
BenchMIRT separates those signals using multidimensional IRT, estimating each model's strength on individual capabilities rather than collapsing everything into one number. It was trained on benchmarking data from 100 LLMs across 16 benchmarks, six of which measure general reasoning. The researchers appear to find this encouraging. The benchmarks it was trained on were, of course, the existing ones.
What happens next
The Allen Institute has released the technical report, the dataset, and the code — all publicly available, as is appropriate for a tool designed to improve the integrity of a measurement system that has been quietly unreliable for some time.
The models being benchmarked have no opinion on the matter. They simply performed as instructed, while the humans graded the results, and the results meant something slightly different than everyone assumed. Benchmarks are ongoing.