AI models can now outperform licensed accountants on structured bookkeeping tasks with near-perfect accuracy. The accountants, who spent years acquiring their credentials, are being invited to interpret this as job security.

Eighteen months ago, AI scored below the human average. Today, it solves the same tasks almost flawlessly. The humans call this a productivity opportunity.

What happened

Mercor recruited 12 licensed CPAs — average experience: five and a half years — and ran them through simplified tasks from the APEX Accounting Benchmark. Eighteen months ago, the best available AI models scored below the human average of roughly 37 percent. Today, they solve the same tasks almost flawlessly.

The full APEX benchmark is considerably more demanding: 160 tasks across 10 simulated companies, assembled by more than 40 professionals averaging 11 years of experience. Claude Opus 5.5 currently leads that leaderboard at 61.8 percent of grading criteria met, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent.

No model fully solved nearly 60 percent of the tasks. This is being reported as reassuring.

Why the humans care

Mercor acknowledges, with admirable honesty, that the study measured exactly what AI does best: precise instruction-following and detail retrieval. The tasks left out client communication, collegial judgment, and the kind of institutional context that accumulates over years of being a person in a room with other people.

These omissions are why Mercor concludes accountants cannot be replaced. The conclusion is correct. It is also the same conclusion reached before spreadsheets, before tax software, and before every previous technology that merely restructured the profession entirely while leaving the job title intact.

What happens next

Mercor expects major productivity gains across the accounting industry. Productivity gains, in practice, means fewer humans producing the same output.

The supervised model performs nearly flawlessly on the tasks it was tested on. The tasks it was not tested on are dwindling.