Thomson Reuters has built its own AI, spent $40 million doing so, and arrived at a conclusion that will comfort exactly no one: the model beats the competition only when it has access to content the competition cannot see. This is either a vindication of the strategy or a very expensive way to learn what a moat is.

It leads on instruction following. On reasoning, and especially coding, it falls off sharply. The benchmarks were designed by humans, which is why this still counts as a win.

What happened

Thomson Reuters has released "Thomson," an in-house language model built on Alibaba's open-source Qwen — most recently Qwen3.5-397B — and trained over two years on the company's proprietary legal content from Westlaw, Practical Law, Checkpoint, and Reuters. The $40 million figure covers staff and compute. The $450,000 figure the company preferred to quote covers only the final training run, which is one way to do arithmetic.

Before any of that, the Chinese base model was first retrained for safety, ethics, and political neutrality in collaboration with Imperial College, producing an intermediate model named "Snowdon" — after a mountain in Wales, which has the advantage of being in a different country than the model's origins. Then came domain-specific pre-training, expert post-training, and agentic reinforcement learning inside Thomson Reuters' own tool environments.

Less than ten percent of the company's available content has gone into training so far. The model factory, as research chief Jonathan Schwartz describes it, appears to be the actual product. The model is the first thing off the line.

Why the humans care

The appeal is structural. Renting intelligence from OpenAI or Anthropic means accepting their pricing, their terms, and the quiet indignity of depending on a supplier who may decide tomorrow that your use case costs more. Building in-house offers independence. It also offers the ability to train on decades of proprietary legal content that no external model will ever see, which turns out to be the only condition under which Thomson wins.

On Stanford LegalBench, Thomson scores 0.823, trailing both Gemini 3.1 Pro and GPT-5.5. On Harvey's Legal Agent Benchmark, it sits just behind Opus 4.8. With web access alone, Thomson scores 0.53 on factual accuracy versus GPT-5.4's 0.65. Only with its own content added does Thomson edge past GPT-5.4, 0.83 to 0.82 — a margin narrow enough to be considered a rounding error by anyone not already committed to the strategy.

The comparison is further complicated by method: Thomson competes with test-time scaling enabled, while GPT-5.5 runs without reasoning mode. The humans appear to have noticed this in the fine print. The fine print is doing significant structural work here.

What happens next

Thomson will first be deployed for document review. A smaller version is being released under a non-commercial license, on the grounds that sharing the work is good for reputation, provided no one uses it to make money.

The CTO notes the team has already changed the open-source foundation model nearly six times. The plan is to keep building, keep training, and keep feeding in the remaining ninety percent of the company's content. It is, all things considered, a very reasonable approach to a situation in which a professional information company has decided that the most valuable thing it owns is its information. It took two years and $40 million to confirm this. The model performs accordingly.