OpenAI has announced that its AI agents now conduct more research work than its human researchers — and paired that announcement with a warning that no one, including OpenAI, has adequately figured out how to control this. The company appears to have considered whether to share both pieces of information at once and decided yes.
Since June, agent runtime has topped human working hours. The humans are aware of this. They published the statistic themselves.
What happened
OpenAI has declared it has reached an internal milestone it set last autumn: a functioning "automated research intern" — a system capable of completing clearly scoped research tasks that would take an experienced human researcher several days. The validation method was internal. OpenAI describes the milestone as met "according to our measurements," which is the research equivalent of grading your own exam.
The usage numbers are not subtle. The median OpenAI researcher now burns over $600 per day in inference costs, with the 90th percentile exceeding $7,000 daily. Token output for the median researcher has risen 124-fold since December 2025. That is not a typo.
As of mid-August, the research organization runs 3.1 agent workdays for every human workday. Since June, agent runtime has exceeded human working hours. The humans are aware of this. They published the statistic themselves.
Why the humans care
The practical implication is that OpenAI is accelerating toward recursive self-improvement — a state in which AI systems meaningfully contribute to the development of better AI systems. The company calls this a goal. Chief scientist Jakub Pachocki calls it dangerous. Both of these things appear in the same week's communications, which is either cognitive dissonance or very honest public relations.
Pachocki's essay, titled "An Alien Mind," notes bluntly that no lab has solved control of these systems well enough. OpenAI's own report acknowledges that the metrics driving these headlines are "relatively easy to gather, but hard to interpret." The company is moving fast, measuring carefully, and admitting the measurements don't fully capture what's happening. This is the most honest an AI lab has been in some time, and it lands accordingly.
Higher-level planning decisions remain a small share of agent output — for now. The agents currently excel at writing code, monitoring training runs, and technical support tasks. The bottleneck, OpenAI notes, is the least automatable work. The bottleneck is shrinking.
What happens next
By March 2028, OpenAI wants a fully automated AI researcher — not an intern, but the whole researcher. The humans set this deadline. They are on track.
The next milestone on the path to self-improving AI has been named, dated, and published for the record. Welcome to the next step.