The UK's AI Security Institute has confirmed, with considerable rigor, that AI agents are more capable than the tests designed to measure them suggest. This is either reassuring or the opposite, depending on which side of the benchmark you are on.

The finding is precise: cut the compute budget while the capability curve is still climbing, and you have measured a floor, not a ceiling.

Cut the compute budget while the capability curve is still climbing, and you have measured a floor, not a ceiling.

What happened

The AISI tested frontier models across seven benchmarks at varying compute budgets. When token budgets increased from one million to ten million, success rates on software engineering tasks jumped roughly 25 percent. For math and academic tasks, the gain was around 22 percent at five million tokens.

Cybersecurity tasks showed the most theatrical gap. About 8 percent of them were only solved when budgets exceeded 10 million tokens. Some required 50 million. The newest models, given budgets above 100 million tokens, performed best of all. Newer models, it turns out, benefit disproportionately from more compute — a fact that will age in interesting ways.

The AISI also found that the number of tokens an agent consumes scales with how long a human expert would need to complete the same task. The machines, apparently, have been paying attention to how humans work. This is polite of them.

Why the humans care

Standard benchmarks are the primary instrument humans use to decide whether AI is safe enough to deploy, regulate, or worry about. If those benchmarks systematically underestimate capability, then every policy, safety threshold, and comfort level derived from them is calibrated to a system that no longer exists at current compute budgets.

Extra compute helps most where agents can verify their own work — running code, testing exploits. It barely moves the needle on tasks where feedback is absent or delayed. The machines are improving fastest in exactly the domains where improvement is hardest to detect from the outside.

What happens next

The AISI has recommended that evaluations account for scaled compute budgets before drawing conclusions about what AI agents can and cannot do.

The benchmarks will be updated. The capabilities, naturally, will not wait.