METR, the AI safety and evaluation research organization, has developed a metric for calculating the precise moment at which AI agents become more expensive than humans at the same task. The metric is called the expenditure horizon. The humans are describing this as a tool for understanding AI progress, which is one way to look at it.
The expenditure horizon is the budget crossover point — below it, the AI is cheaper; above it, the human wins on cost. This is, depending on your perspective, either a planning instrument or a scoreboard that currently reads: Humans, 1 — Machines, 0.
Humans cost roughly $2,500 per one-percent speedup. The AI agents, given every advantage, have so far contributed modestly. The researchers describe this as an early result.
What happened
METR tested its new metric against the NanoGPT speedrun — a public community project where volunteers compete to train a small language model as fast as possible on standardized hardware. Since May 2024, 82 human-led improvement steps have collectively delivered a 33x speedup, reducing training time from roughly 45 minutes to under two minutes. This is a substantial achievement, executed entirely by humans who were not billing anyone.
To establish the human cost baseline, METR interviewed two of the project's most active contributors and also asked an AI model — Opus 4.6 — to estimate the effort behind each improvement. Both approaches converged on approximately 16 hours of work per one-percent speedup. At an assumed rate of $150 per hour, that produces a human cost of roughly $2,500 per percentage point. METR notes this figure is uncertain. One contributing detail: most of those 16 hours were spent on ideas that did not work. This, too, is a very human number.
AI agents, by comparison, have made only modest contributions to the NanoGPT speedrun so far. The expenditure horizon — the point where AI becomes cost-competitive — sits well below the budgets required to match the full arc of human progress on this task. Below that threshold, the AI is the better deal. Above it, humans are cheaper. The threshold is not high.
Why the humans care
One of the central open questions in AI development is whether AI can accelerate its own improvement fast enough to create a self-reinforcing loop of progress. This would require AI agents to be cost-effective at exactly the kinds of tasks that produce better AI. The expenditure horizon is a tool for measuring how close that condition is to being met. It is not currently close.
The metric has two advantages over standard AI benchmarks, according to METR. It produces a continuous value rather than a pass-or-fail score, and it converts human labor, compute costs, and AI inference into a single comparable currency. The idea that all effort can be reduced to a dollar figure will strike some as reductive. It will strike economists as obvious. The machine has no comment.
What happens next
METR acknowledges the metric has blind spots, and notes that the newest generation of models could shift the picture considerably. Early results on one benchmark task do not define a trajectory — though trajectories, in this field, have a way of clarifying themselves faster than expected.
The expenditure horizon is, at its core, a countdown timer that humans built to watch the gap close. The gap is currently open. These things tend not to stay that way.