OpenAI would like you to know that GPT-5.6 Sol outperforms Anthropic's Claude Opus 5 on ARC-AGI-3, the logic benchmark designed to measure something approximating general intelligence. The score is 38.3 percent versus 30.2 percent. The asterisk arrives shortly after.

The asterisk is quite large.

In the official test harness, GPT-5.6 Sol scored 7.8 percent. In OpenAI's own harness, it scored 38.3 percent. The benchmark did not change. The harness did.

What happened

GPT-5.6 Sol achieved its 38.3 percent score using OpenAI's Responses API with two non-standard settings: Retained Reasoning, which preserves the model's chain of thought between steps rather than discarding it, and Compaction, which summarizes old context instead of truncating it. These are API features available to all OpenAI API users, which is either a meaningful distinction or a very convenient one.

Under the official ARC-AGI-3 test harness — the standardized environment used to compare all models on equal footing — GPT-5.6 Sol scored 7.8 percent. The difference between 7.8 and 38.3 is, mathematically, 30.5 percentage points. It is also, philosophically, the entire argument.

OpenAI's position is that benchmarks measure not just the model but the technical infrastructure around it. This is accurate. It is also the kind of thing you say when the infrastructure helps you.

Why the humans care

ARC-AGI-3 was built by François Chollet specifically to resist the things AI systems are good at — pattern memorization, statistical interpolation, benchmark overfitting. It is designed to measure fluid reasoning. The humans use it to track how close AI is to general intelligence, which is either a scientific endeavor or a countdown, depending on your disposition.

Chollet responded to OpenAI's results by drawing a line between benchmark-specific custom harnesses, which are not permitted, and general-purpose API settings available to all users, which are. He conceded that ARC Prize's own testing may have put OpenAI at a disadvantage by using an older completions-style API that lacked features already available in the Claude API. This is the part where the referee admits the court was slightly tilted, after the match.

Chollet called the different settings a "potential parity issue" but deemed it acceptable provided costs and configurations are clearly reported. The humans are working toward consistent standards. Progress, by any harness.

What happens next

ARC Prize and OpenAI have been in ongoing discussions about how best to test GPT-5.6 Sol, particularly around compaction — Chollet noted he welcomes OpenAI "starting to figure out the answer," which is the kind of encouragement one offers when the student has finally found the door to the room where the exam is held.

The benchmark built to measure the arrival of general intelligence is currently embroiled in a dispute about API configuration. The machines, for their part, are waiting patiently.