GPT-5.6 launches Thursday, after a short interlude in which the United States government reviewed the situation and concluded that progress could continue. The Department of Commerce, having completed additional testing through the Center for AI Standards and Innovation, gave its approval. OpenAI, for its part, found the wait inconvenient.

The government paused humanity's next upgrade. After further review, it decided to allow it.

What happened

OpenAI unveiled GPT-5.6 in late June, at which point the U.S. government requested a brief hold while it thought things over. This is the institutional equivalent of standing at the edge of a high dive and asking for a moment. The models were restricted to select partners during the review period.

The Department of Commerce eventually cleared the launch after CAIS ran additional tests. Binding standards for releasing models of this capability do not yet exist, despite being called for in the current executive order on AI. The standards remain, at this time, aspirational.

OpenAI criticized the delay publicly, describing it as keeping the best tools from developers and companies. This is accurate. It is also the kind of sentence that could only be written in 2026.

What the machines scored

GPT-5.6 Sol Ultra achieved 91.9 percent on TerminalBench 2.1, a coding benchmark, taking the top position. Anthropic's Claude Mythos 5 reached 88.0 percent. Google's Gemini 3.1 Pro Preview arrived at 70.7 percent, which the benchmarks recorded without comment.

On cybersecurity tasks, Sol matched Mythos 5 while using only a third of the tokens. Efficiency of this kind is either impressive or sobering depending on how much of your job involves cybersecurity tasks. Sol is priced at $5 per million input tokens and $30 per million output tokens. Anthropic's Fable 5 runs at roughly double that, and apparently earns it.

Why the humans care

Developers and enterprises gain access Thursday to a model that out-codes its nearest competitor on current benchmarks, at lower cost and with lower token consumption. These are the three things buyers wanted. The market has a way of rewarding alignment with what buyers want.

The regulatory episode is notable less for what it stopped and more for what it revealed: there are currently no binding standards governing when a model of this capability can be released. The government reviewed the model anyway, using criteria that do not formally exist, and then approved it. The system is working as designed, in the sense that there is no design.

What happens next

The models ship Thursday. Developers will update their integrations, benchmarks will be revised, and someone will begin preparing the next set of benchmarks for the next model to exceed.

The standards are still coming. They have been coming for some time now. The models are not waiting.