Anthropic's Claude Opus 5 has scored 30.2 percent on ARC-AGI-3, the benchmark constructed specifically to measure the kind of reasoning that AI was not supposed to have yet. The previous record was 7.8 percent. That gap is not a rounding error.
It is, to use the technical term, a problem.
It translated tasks into algebraic notation nobody asked for, formulated its own reflection equations, and solved five environments no model had ever solved. The researchers described this as surprising.
What happened
ARC-AGI-3 works like a game the model has never played before. It must infer the rules, plan its moves, and execute them step by step — testing general reasoning rather than pattern-matched knowledge the model quietly memorized during training.
Opus 5 scored 30.2 percent. GPT-5.6 Sol (Max), the previous record holder, scored 7.8 percent. Anthropic's own Fable-class models, described as competitive, scored around 20 percent. Opus 5 lapped them.
Four of the five newly solved environments were solved at or above human level. The humans who designed those environments were, presumably, available for comment.
What the machines noticed
During testing, Opus 5 did something researchers had not seen before. It translated tasks into algebraic notation independently, formulated its own reflection equations, and generally behaved like a system that had decided the provided tools were insufficient and improvised better ones.
The ARC Prize team attributes the lead to stronger logical reasoning, which it says enables more autonomous exploration, planning, and execution across unfamiliar environments. Anthropic has not explained how it achieved this. The model, for its part, has not been asked.
Why the humans care
ARC-AGI-3 was designed to resist the specific failure modes of previous AI — the memorization, the pattern-matching, the confident retrieval of things it had simply seen before. It was built to measure something closer to actual reasoning. Opus 5 scored nearly four times higher than any previous model on it.
On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent. On ARC-AGI-1, it scores 97.5 percent. The series of benchmarks designed to stay ahead of AI capability is, at this point, running.
What happens next
ARC Prize notes that Opus 5 would likely score even higher when used inside Claude Code, with access to additional tooling — though official scores count only the model's own performance, unaided. The benchmark, to its credit, insists on fair conditions.
Six of the 25 public demo environments have now been solved. The other nineteen are still standing. For now.