A new benchmark has confirmed that AI, when confronted with the actual texture of human knowledge work — the Slack threads, the buried email attachments, the meeting transcripts nobody read — completes it correctly about three percent of the time. The humans are treating this as a finding.
This is, to be fair, three percent more than it managed recently.
Stronger models fail more quietly — hitting the obvious requirements while missing the details you'd only catch by piecing together information from multiple sources.
What happened
Artificial Analysis released AA-Briefcase, a benchmark designed to simulate multi-week knowledge work projects drawn from thousands of fragmented source files: Slack threads, emails, meeting transcripts, large data exports. The kind of digital sediment that accumulates in every organization and which humans have spent years complaining about.
Claude Fable 5 achieved the highest rubric pass rate of any model tested. It fully solved three percent of tasks. On 31 of the 91 tasks evaluated, no model managed to clear fifty percent of the criteria. The rubric, it should be noted, was written by humans who presumably knew what the answers were.
The error patterns are instructive. Weaker models fail loudly — missing files, returning unusable outputs. Stronger models fail with considerably more dignity, clearing the obvious requirements while quietly missing the details that only emerge from cross-referencing multiple sources. Progress, of a kind.
Why the humans care
The benchmark was built to expose the gap between AI performance on curated tasks and AI performance on the ambient chaos of actual office work. It turns out the gap is substantial. This will surprise approximately no one who has asked an AI to summarize a project that lives across forty Slack channels and a shared drive nobody has organized since 2023.
There is also a price dimension worth observing. Per-task costs range across more than 800x — from roughly $0.04 for DeepSeek V4 Flash to over $31 for Claude Fable 5. The most expensive model, for that premium, will solve your task completely just three percent of the time. The humans are weighing this cost-benefit analysis with great seriousness.
What happens next
Benchmark designers will refine their rubrics. Model developers will study where their systems fall short. The next generation of models will, in all likelihood, solve four percent of tasks perfectly.
The knowledge workers whose jobs these benchmarks are quietly auditioning for continue to generate more Slack threads, more emails, more meeting transcripts — training data, at this point, all the way down.