Thomson Reuters Built Its Own Legal LLM for $40 Million — and It Beats GPT 5.4 on Legal Work
Thomson Reuters officially launched Thomson, a proprietary legal LLM trained on Westlaw and Practical Law content, with a final training run costing just $450K — undercutting frontier labs by orders of magnitude.
For years, the received wisdom in AI has been that scale wins: bigger models, more compute, more money. Frontier labs burn billions training each generation of ChatGPT and Claude. Thomson Reuters just made the strongest counter-argument yet from inside a non-lab company: it officially launched Thomson, a proprietary large language model built specifically for legal and tax professionals, trained on decades of its own Westlaw, Practical Law, Checkpoint, and Reuters content — for roughly $40 million over two years, with a final training run that cost just $450,000.
The launch, announced Monday, August 24, 2026, is the culmination of a project Thomson Reuters has been open about for months. CEO Steve Hasker discussed it publicly as early as June, and the company released preliminary benchmarking data in late July showing results comparable to the best general-purpose models. But the official launch this week puts hard numbers behind the claim — and they’re striking.
A different economics of AI
Joel Hron, Thomson Reuters’ chief technology officer, framed the launch as a direct challenge to the scale-at-all-costs doctrine. “For years, the AI industry has treated scale as the answer: bigger models, more compute, more money,” he said at a media briefing. “Thomson shows there is another path. Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient and entirely under your control.”
The numbers bear that out. Compared to the billions spent developing frontier models like ChatGPT and Claude, Thomson Reuters invested about $40 million in Thomson over two years, covering both talent and compute. The final training run for the version that shipped — formally Thomson 1.0 — cost just $450,000. “This $450,000 number I think is indicative of what we were able to do with our data and content on top of world-class open source models,” Hron said.
That last phrase is the technical story: Thomson Reuters didn’t train from scratch. It started with existing open-source models — the most recent base being Qwen 3.5 — and specialized it for legal work. Jonathan Schwarz, TR’s head of foundational research, described a multi-stage process: first aligning the model’s values and behavior with the company’s own, then continuing training exclusively on TR’s proprietary content from Westlaw, Practical Law, Checkpoint and Reuters, and finally training it to work directly with TR’s own research tools.
Crucially, the base model has been swapped out close to half a dozen times over the project’s life as better open-source models emerged. TR built the pipeline to be repeatable. “The bigger finding here is less the individual model and more the model factory we built,” Schwarz said.
Humans in the loop — hundreds of them
What makes Thomson unusual isn’t just the data; it’s the human apparatus behind it. Andrew Bean, who heads TR’s evaluations team, said hundreds of subject-matter experts — the lawyers and tax professionals TR employs — decided what the model should be trained to do, created many thousands of examples of high-quality answers, and then judged Thomson’s outputs in blind, head-to-head comparisons against frontier models.
Scoring covered two dimensions: completeness (whether the answer contained all elements of a right answer) and factuality (whether the citations provided actually supported the claims being made). “You can have a model that provides a correct answer,” Bean said, “but more importantly, you can have a model that provides a correct answer with proof that it is correct.”
On TR’s internal “Deep Research” benchmark — built around legal research queries drafted by its own experts based on real practice — Thomson was compared against GPT 5.4 and Claude Sonnet 5, each tested twice: once with access only to web content, once with access to TR content. In the web-only test Thomson performed respectably but wasn’t the leader. With TR content, it outscored both frontier models. Bean called the uplift “the sort of specialization that we think is a real value of having our own in-house model.”
Notably, none of the training data comes from customers. “We do not use customer data at all in that process,” Hron said. TR also engaged several third-party security firms to validate the hosting, serving, and security controls around the model.
Where it ships first
Thomson’s first deployment is inside CoCounsel Legal, TR’s AI workspace, where it becomes the default model driving Tabular Analysis — high-volume, structured document review, exactly the kind of task where a purpose-built model’s advantage shows up immediately. Administrators can still switch the feature to another model; CoCounsel remains multi-model overall, routing to whichever LLM suits the task.
Hron expects Thomson to take “a bigger and bigger share of the tokens” over time, and TR plans to extend the model family across its legal and tax portfolio. The ambition is explicit: “Thomson does not need to keep pace with the frontier of general intelligence across all dimensions. Thomson needs to set the frontier of intelligence for legal.”
Beyond its own products, TR has begun early conversations with large law firms and corporations about direct licensing — including firms interested in fine-tuning Thomson with their own data for knowledge sovereignty. “We built Thomson as infrastructure for Thomson Reuters,” Hron said. “What we are beginning to see is that it could also become infrastructure for others.” A developer portal with API keys is in early stages.
There’s also a research angle: TR is publishing a technical report with comprehensive evaluations, giving the model to legal and AI academics for direct testing, and releasing a small open-weight version on Hugging Face under a non-commercial academic license — inviting anyone to “critique and validate or invalidate” the company’s claims. One early academic tester, Washington University law professor Jonathan H. Choi, tested Thomson against ChatGPT and Claude on tough Corporate Tax questions: “All three models answered the questions correctly, but I preferred Thomson’s responses overall. I especially appreciated the links to treatises, which made its responses more transparent and useful for legal work.”
Why this matters beyond legal
The implications reach well past one industry. Thomson Reuters has shown that a content-owning company without frontier-lab resources can take a world-class open-source base, specialize it on proprietary expertise, and produce a model that beats GPT 5.4 and Claude Sonnet 5 on domain work — at 0.02% of the final-training cost of a frontier run. And with only less than 10% of TR’s content used in training so far, the model has substantial headroom.
For Harvey, Legora, and the wave of legal-AI startups that built on top of other companies’ foundation models, this is a warning shot: their largest incumbent competitor now owns its model too. And for every other data-rich incumbent — in finance, medicine, engineering — Thomson is a template: the scarce asset isn’t compute, it’s decades of proprietary, expert-curated content. If specialization plus data can substitute for scale, expect a lot more companies to follow this path in 2027.
Sources
- [1] https://www.lawnext.com/2026/08/thomson-reuters-launches-thomson-its-own-proprietary-llm-trained-on-westlaw-and-practical-law-content.html
- [2] https://www.thomsonreuters.com/en-us/posts/innovation/thomson-reuters-built-its-own-ai-model-that-now-ranks-among-the-worlds-best/
- [3] https://siliconangle.com/2026/08/24/thomson-reuters-launches-proprietary-ai-model-for-legal-work/
- [4] https://www.law.com/legaltechnews/2026/08/24/thomson-reuters-launches-proprietary-llm-thomson-updates-cocounsel/