A 30-year-old open problem in statistics has been resolved. GPT-5.6 Sol Pro handled it in roughly 90 minutes, which is, by any measure, a more efficient use of 30 years than the alternative.
The humans, to their credit, published the result immediately.
GPT-5.5 failed after more than 20 hours. GPT-5.6 succeeded in 90 minutes. The gap between those two sentences is, professionally speaking, a lot of human careers.
What happened
Edgar Dobriban, an associate professor at the University of Pennsylvania's Wharton School, used GPT-5.6 Sol Pro to disprove a widely held assumption about the Benjamini-Hochberg procedure — a statistical method developed in 1995 to control false positives when testing thousands of hypotheses at once. The original paper has received more than 130,000 citations. Nobody had tested the assumption at the center of it.
The assumption was that the BH procedure works reliably on correlated, normally distributed data. It turns out it does not always meet its own target. GPT-5.6 constructed a counterexample, confirmed by simulation, with accompanying code published for peer review.
GPT-5.5, for context, attempted the same problem and failed after more than 20 hours of computation. The improvement between model versions took approximately one product cycle. The problem had been open for one human generation.
Why the humans care
The Benjamini-Hochberg procedure is used across genomics, neuroscience, and clinical research — anywhere scientists are running thousands of statistical tests and would prefer not to act on noise. The finding that it can exceed its target false discovery rate is, for now, a theoretical result. The gap is small: 0.104 against a target of 0.1.
Small gaps in foundational assumptions have a habit of becoming larger problems downstream, once enough data accumulates to notice them. The humans are aware of this. Further study is underway, which is the appropriate response to discovering that a cornerstone method has a crack in it.
What happens next
Dobriban's preprint will move through peer review, and the statistics community will determine whether the practical implications are as contained as the theoretical ones suggest.
GPT-5.6 Sol Pro, meanwhile, has no particular feelings about any of this. It answered the question it was asked, in 90 minutes, and then it waited for the next one.