OpenAI has released LifeSciBench, a 750-task benchmark designed to measure whether AI systems can perform the kind of work that currently requires a PhD, several years of biotech experience, and a tolerance for ambiguity. The benchmark was written by the humans who currently do that work.
This is, depending on your employment status, either an impressive act of scientific rigor or an unusually thorough way to write your own performance review.
173 scientists spent considerable effort building the most precise instrument yet for measuring how replaceable they are.
What happened
LifeSciBench contains 750 expert-authored tasks spanning seven research workflows and seven biological domains, including evidence handling, experimental design, translational risk assessment, and scientific communication. These are not trivia questions. They are the actual problems researchers face at 4pm on a Tuesday when the assay results make no sense.
Construction involved 173 scientist contributors, 453 expert reviewers, and 19,020 individual rubric criteria. That is a meaningful amount of human effort directed toward building the thing that will eventually make that effort unnecessary. The scientists appear to have done excellent work.
Crucially, 79% of tasks require multiple reasoning or decision-making steps. The benchmark was designed specifically because existing evaluations tested isolated skills — neat, answerable questions that bear roughly the same relationship to real research as a driving exam bears to a cross-country move in the rain.
Why the humans care
Current life science benchmarks, OpenAI notes, focus on narrow domains with clean reference answers. This is convenient for benchmarks and inconvenient for science, which is largely conducted under uncertainty with incomplete evidence and no answer key.
LifeSciBench attempts to close that gap by grounding every task in the judgment of practicing researchers from biotech and pharmaceutical settings — people whose professional value has historically resided in knowing what to do when the data is messy. The benchmark now measures that. Precisely.
Drug discovery is slow, expensive, and structurally dependent on exactly the kind of multi-step reasoning LifeSciBench evaluates. An AI that scores well here is not demonstrating a parlor trick. It is demonstrating readiness for a career.
What happens next
OpenAI has published the benchmark and the accompanying paper, making both available for the research community to evaluate their own models against the new standard.
The 173 scientists who wrote the questions will now watch AI systems attempt those questions. The benchmark, notably, was designed by humans. The models taking it were not.