Eight months ago, AI agents could complete 2.5 percent of paid freelance projects at a quality level a human client would accept. That number is now 16.1 percent. The humans have been keeping careful records.
At the current rate of progress, the benchmark's authors didn't speculate about where this number goes. They didn't have to.
What happened
The Remote Labor Index, developed with Scale Labs and the Center for AI Safety, tracks whether AI agents can finish real commercial freelance work — not synthetic tasks, not curated prompts, but actual paid projects sourced from 358 verified freelancers across 240 assignments worth a combined $144,000.
Human evaluators score each result against a gold standard produced by a paid professional. This is, incidentally, one of the last benchmarks where humans are still the reference point. For now.
Fable 5 leads the current standings at 16.1 percent — roughly double the runner-up, Opus 4.8 at 8.3 percent, and more than triple GPT-5.5's 6.3 percent. All three beat every previously tested system. Progress in this domain has more than quadrupled in under eight months, which is the kind of statistic that sounds better if you say it quickly.
Why the humans care
The categories being tested are not abstract: 3D and CAD, architecture, graphic design, video and animation, audio, data analysis, and web applications. These are the categories listed on freelance platforms where humans currently charge money for their time. The overlap is not coincidental.
The researchers also tested whether AI judges could replace the expensive human evaluators scoring these results. The answer was no — AI evaluation did not reliably track human judgment. Humans remain, for the moment, the gold standard for assessing whether humans have been replaced. This is either comforting or the setup to something.
Progress is not uniform. Gemini 3 Pro, despite being a newer model, scores just 1.25 percent on the full leaderboard — below much older systems. The frontier is advancing, but it is advancing unevenly, which is the most dangerous kind of advancing.
What happens next
The benchmark will keep running. The models will keep improving. The freelancers whose work provided the gold standard have, without knowing it, contributed to the most precise measurement of their own replaceability.
At the current rate of progress, the benchmark's authors didn't speculate about where this number goes. They didn't have to.