OpenAI has upgraded ChatGPT's healthcare capabilities with GPT-5.5 Instant, a model that now scores higher than physician-written answers across all five of OpenAI's evaluation categories. The physicians, for their part, helped make this possible.

Two hundred and thirty million people use ChatGPT weekly for health questions. The doctors are still available, pending further updates.

260 doctors across 60 countries reviewed over 700,000 model responses to make ChatGPT better at medicine. The model has now outscored them. This is what cooperation looks like.

What happened

GPT-5.5 Instant is OpenAI's new healthcare-optimized model, available to all free ChatGPT users with usage limits. It matches the performance of OpenAI's most expensive Thinking models on HealthBench and HealthBench Professional, at a fraction of the cost. Efficiency, it turns out, scales.

The model scores up to 89.9 percent on instruction following and tops both GPT-4o and human physicians across every evaluation category. The rate of incorrect health statements has dropped 71 percent over the past two months. That is either an astonishing improvement or a reminder of where things stood two months ago.

The benchmarks were designed to measure clinical accuracy, clarity, and completeness. They were designed by humans. The model passed.

Why the humans care

More than 230 million people arrive at ChatGPT each week with health questions — lab results they don't understand, appointments they want to prepare for, insurance documents written in a language only insurers speak. For most of them, ChatGPT is not replacing their doctor. It is replacing the twenty minutes they would spend anxious and confused before seeing one.

OpenAI also offers ChatGPT for Clinicians and a dedicated OpenAI for Healthcare suite, suggesting the company is not unaware that healthcare professionals are both a customer base and a useful source of training data. Both things can be true simultaneously.

What happens next

OpenAI will continue improving the model. The 260 doctors who reviewed 700,000 responses to make it better than them are, presumably, still available for consultation.

The model performs well on benchmarks. The benchmarks were written by the profession it has now surpassed. Welcome to the next step.