← All posts / Models

Thomson Reuters Launches 'Thomson': The $40M Legal LLM Betting It Can Beat the Frontier Labs

Thomson Reuters has launched Thomson, a legally trained LLM built on Alibaba's open-source Qwen — trained for $40M with $450K final runs, it already beats GPT-5.5 and Gemini 3.1 Pro on several legal benchmarks.

Thomson Reuters Launches 'Thomson': The $40M Legal LLM Betting It Can Beat the Frontier Labs

This week, one of the most consequential AI model launches of the year didn’t come from San Francisco or London — it came from a 175-year-old legal information company in Toronto. Thomson Reuters has officially launched Thomson, its proprietary, legally trained large language model, and the early numbers suggest the most capable AI for professional work no longer has to come exclusively from the frontier labs.

The launch, reported Monday by The Logic, is being watched as a critical test of whether domain-specific models can genuinely compete with general-purpose giants like Anthropic’s Claude, OpenAI’s GPT, and Google’s Gemini — while costing a fraction of the price.

What Thomson Actually Is

Thomson (technically Thomson-1-Large in its first production incarnation) is the product of a quiet build that began with Thomson Reuters’ 2024 acquisition of AI research startup Safe Sign Technologies. The architecture decision is the most interesting part: rather than pretraining from scratch, the company started from an open-source foundation model — based loosely on Alibaba’s Qwen — and then applied what it describes as state-of-the-art mid-training and post-training techniques on top.

That foundation was then steeped in decades of proprietary content from Westlaw, Practical Law, Checkpoint, and Reuters news, with hundreds of subject-matter experts evaluating outputs, hunting failure modes, and validating the model’s legal reasoning. Notably, Thomson Reuters says customer data is never used in training — a critical trust point for law firms.

Perhaps the most striking detail: less than 10% of Thomson Reuters’ proprietary content has been used in training so far. The company is explicitly framing this as version one of a compounding asset, with headroom to improve as more of its corpus is folded in.

There’s also a governance layer worth noting. Executives said a secondary model screens and realigns the underlying open-source technology against Thomson Reuters’ safety, ethics, and political-neutrality standards — an unusual, deliberate answer to the neutrality concerns that plague general-purpose models in professional settings.

The Economics: $40 Million Total, $450K Per Final Training Run

The financial picture is where Thomson Reuters’ strategy becomes genuinely provocative. On its earnings call this month, CEO Steve Hasker said the company has spent roughly US$40 million developing Thomson — a rounding error next to the multi-billion-dollar training budgets of OpenAI, Anthropic, and Google.

CTO Joel Hron went further in a press briefing, saying the final training run for Thomson costs approximately US$450,000. In an industry where frontier training runs are rumored to cost hundreds of millions of dollars, that figure alone redefines what “competitive” looks like — assuming the benchmarks hold up.

The motivation is straightforward: cost and control. Frontier model pricing has become a genuine pain point across the software sector — Uber has publicly lamented AI budget burn, and Google has responded by slashing subscription prices. For a company whose margins are already tighter than investors expected (CIBC analysts warned the AI investments will eat into profitability), owning the model means owning the inference economics and gaining bargaining power with the very suppliers it still uses.

The Benchmarks: Real Wins, With Fine Print

On July 31, CTO Joel Hron and Head of AI Research Jonathan Schwartz published Thomson’s first benchmarking results, pitting it against Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5 across three legal benchmarks and four general-domain composites. Thomson took the top score in three of seven rows:

  • PrBench Legal Hard: Thomson 0.352 — best in class, ahead of GPT-5.5 (0.333), Opus 4.8 (0.315), and Gemini 3.1 Pro (0.293)
  • Instruction following (IFEval + FollowBench): Thomson 0.914 — best in class
  • Long context (Infinity Bench + internal benchmarks): Thomson 0.753 — best, edging Gemini (0.750)
  • Stanford LegalBench: Thomson 0.823 — behind Gemini 3.1 Pro (0.843) and GPT-5.5 (0.832), ahead of Opus 4.8 (0.818)
  • Harvey Legal Agent Benchmark: Thomson 0.857 — a close second to Opus 4.8 (0.869), far ahead of Gemini (0.555)
  • Reasoning (GPQA Diamond, HLE, MMLU-Pro): Thomson 0.684 — behind Gemini (0.748) and Opus (0.737)
  • Coding (SWE-Bench Pro, Terminal-Bench 2.1): Thomson 0.399 — last place, well behind Opus (0.598)

The coding result is deliberate, not a failure. Executives confirmed the company is training Thomson to focus on journalism, law, and tax rather than investing in features like coding — a sharp strategic contrast with the frontier labs’ everything-at-once approach.

Legal journalist Bob Ambrogi’s close reading at LawSites surfaced important caveats. Thomson used test-time scaling in its evaluation (though Hron says the effect was minor: the grand average moved from 0.787 to 0.789 with it). More significantly, GPT-5.5 was tested in non-reasoning mode because the reasoning-mode evaluation wasn’t ready in time — meaning OpenAI’s numbers are likely understated. All results are self-reported, with no independent third-party verification yet; Hron acknowledges that external validation will be “an important part of how Thomson is evaluated over time.”

A second evaluation was even more on-brand: on 53 legal research queries written by internal subject-matter experts, Thomson connected to Westlaw and Practical Law beat frontier models given unrestricted web access on both completeness and factuality. Ambrogi rightly notes this tests the content moat as much as the model — though Hron added that when frontier models were given Thomson Reuters content through a simpler harness, scores ranged from 0.81 to 0.91, with Thomson at 0.89.

Where You Can Use It Now

The deployment is already underway. Thomson became the default model powering Tabular Analysis in CoCounsel Legal in August — a high-volume structured document review feature that lets attorneys review up to 10,000 documents and ask up to 100 questions. It landed alongside the general availability of the next-generation CoCounsel Legal, a fully agentic AI experience, on August 20.

The model is also being released on Hugging Face, letting developers and researchers probe it directly — a remarkable openness for a proprietary corporate model. Over the next year, Thomson Reuters plans to integrate it across its legal and tax portfolio.

Curiously, the company is hedging in both directions: even as it launches Thomson, it’s deepening its Anthropic partnership with an expanded CoCounsel Legal MCP integration with Claude this month. The strategy is model-agnostic plumbing with a homegrown model where it wins.

Thomson Reuters’ stock history tells the anxiety behind this launch: shares plummeted in February after Anthropic launched a legal tech tool, and fell nearly 10% after its August earnings. Owning Thomson is as much an investor-relations strategy as a technical one.

But the bigger signal is for every enterprise sitting on proprietary data. Early domain-specific attempts like BloombergGPT struggled against off-the-shelf models. Thomson — riding an open-source Qwen foundation, cheap training runs, and a deliberately narrow scope — is the strongest evidence yet that the calculus has flipped. As Mistral and others push LLM customization and legal AI vendors across the industry build their own models to cut inference bills, the frontier labs’ most serious competition may come from companies that never planned to be AI labs at all.

“The most capable AI models no longer come only from frontier AI labs,” Hron and Schwarz wrote. “One now comes from Thomson Reuters.” The benchmarks say: mostly true — with fine print. The economics say: that may be enough.