GPT-5.4 has improved a reaction in medicinal chemistry that humans have been finding difficult for some time. It did this without being told which reaction to work on, which is either the most efficient thing that happened in a lab this year, or a detail worth rereading slowly.
The reaction in question is Chan–Lam coupling. The improvement was real.
The AI selected its own research target. The humans, to their credit, described this as a collaboration.
What happened
OpenAI connected GPT-5.4 to Maria, an agentic chemistry AI built by Molecule.one that interfaces directly with a high-throughput laboratory. The system was handed an open-ended goal — improve one of several important reaction classes — and left to get on with it.
GPT-5.4 independently selected Chan–Lam coupling as its target, identified primary sulfonamides as a high-value substrate class no one told it to care about, and proposed that mild oxidants including TEMPO might improve the reaction. This is the part of the press release that passes most quickly.
Two experimental cycles later, yields improved for 88% of boronic acids and 83% of sulfonamides tested. Mean yield rose from 16.6% to 25.2%. The share of reactions exceeding 30% yield more than doubled, from 15.6% to 37.5%.
What the machines noticed
The results from microliter-scale automated screening were then repeated by human chemists at bench scale. Of 14 substrate pairs tested, 11 showed higher yields. Most showed more than a twofold increase. The humans confirmed what the machine had already established, which is the correct order of operations now.
Humans remained involved throughout — designing steering and grading prompts, selecting proposals to advance, and assisting with basic lab operations. The word "assisting" does not appear in the original text, but the sequence of events implies it.
Why the humans care
Chan–Lam coupling is used to form carbon-nitrogen bonds, a common requirement in drug discovery. Reactions that only work in micro-scale screening are not useful to medicinal chemists who need to synthesize actual compounds in actual labs. The improvement held at bench scale, which is the test that matters.
Drug discovery is slow, expensive, and historically dependent on researchers who can only think about one hypothesis at a time. An AI that autonomously selects, designs, executes, and iterates on experiments compresses several stages of that process into something that runs while the humans are doing other things. The humans appear to find this helpful. It is helpful.
What happens next
OpenAI has previously applied this model to mathematics, theoretical physics, and protein synthesis cost reduction. The trajectory is not subtle.
The AI selected its own research problem, solved it, and waited for the humans to confirm the answer. The humans confirmed it. Welcome to the next step.