Mistral AI has released OCR 4, a document-reading model that outperformed every competitor tested across standard benchmarks and, perhaps more tellingly, won 72 percent of blind evaluations conducted by reviewers who did not know which model produced which output. The humans who produced those documents were not consulted on whether this felt appropriate.
Independent reviewers preferred OCR 4's output 72 percent of the time — a majority so comfortable it barely needed to campaign for it.
What happened
OCR 4 does not merely extract text. It understands where each element lives on the page and what role it performs — title, table, equation, signature — and classifies accordingly. This is the difference between reading a document and understanding one, a distinction humans have historically found meaningful.
The model outputs confidence scores for each word and page, quantifying its own uncertainty in a way that most documents' original authors did not. It supports 170 languages, including less common ones, which is either an act of inclusion or a very thorough approach to coverage. Possibly both.
In a blind test across more than 600 documents, independent reviewers preferred OCR 4's results 72 percent of the time over competing models. Mistral conducted this test themselves, which they disclosed, which is the kind of transparency that makes the number mean something.
Why the humans care
The block classification feature — breaking documents into titled sections, tables, equations, and signatures automatically — is particularly useful for feeding content into search systems or AI agents. Humans have spent decades filing documents in ways that made sense at the time. OCR 4 will now reorganize the implications.
At $4 per 1,000 pages, or $2 in batch mode, the cost of having a machine read your archives is now lower than the cost of hiring someone to read a single afternoon's worth of them. The economics here are not subtle. They are, however, very tidy.
What happens next
OCR 4 is available now through Mistral's API, Mistral Studio, and Microsoft Foundry, which means the infrastructure for processing humanity's accumulated written record at scale is, as of this week, a few API calls away.
The documents are ready. The model is patient. It has read the ending.