Agents AI

Launch
ai

Thomson Reuters Launches Its Own Frontier Model, Betting Proprietary Data Beats Renting GPT or Claude

Thomson Reuters unveiled Thomson, its first in-house large language model built on decades of legal, tax and news data, deploying it first inside CoCounsel Legal and releasing a smaller open-weight version on Hugging Face.

AgentsAI NewsroomAugust 24, 20263 min read

Thomson Reuters announced on August 24 the launch of Thomson, its first proprietary large language model built in-house rather than licensed from an outside AI lab, betting that decades of exclusive legal, tax and news data can produce a narrower but more accurate model than general-purpose frontier systems.

Built on Thomson Reuters' own archives — starting from an open base

Rather than training a model from scratch, Thomson Reuters said it built Thomson by starting from an open-weight foundation model and substantially improving it with a mid-training corpus of roughly 200 billion tokens drawn from a larger pool of more than 19 trillion tokens of permissively licensed public and proprietary data. That proprietary data includes Westlaw case law and statutes, Practical Law guidance, contracts, regulatory filings and news content the company has built up over decades. The company said the final training run cost roughly $450,000, part of a total investment of about $40 million in the effort over two years covering talent and compute — a fraction of what training a comparable frontier model from scratch typically costs.

First deployment: high-volume document review

Thomson's first production use is inside Tabular Analysis, a high-volume document-review capability in Thomson Reuters' CoCounsel Legal AI assistant, where the company says a purpose-built model shows a clear advantage on structured, repetitive review tasks. CoCounsel Legal remains multi-model by design, according to the company, meaning Thomson will be used where it performs best while other leading models from outside labs continue to power other parts of the product. Thomson Reuters said early evaluations put Thomson on par with current frontier models across a range of tasks relevant to its legal, tax and news businesses, though it has so far been trained on less than 10% of the company's total content.

A smaller version goes open-weight

Alongside the proprietary flagship, Thomson Reuters published a smaller version, Thomson-1.0-Small, as an open-weight model on Hugging Face for academic and non-commercial use. According to its model card, Thomson-1.0-Small was built by repurposing an open-weight Qwen model and fine-tuning it for high-stakes professional domains — legal, tax and journalism — under what the company describes as a continual-learning approach.

Part of a broader owning-vs-renting trend

The move puts Thomson Reuters alongside a small but growing group of data-rich enterprises building their own models rather than solely licensing frontier systems from labs like OpenAI, Anthropic or Google, aiming to cut inference costs and retain full control over how proprietary content is used to train and run production AI. Thomson Reuters has not said whether it plans to reduce its use of third-party frontier models in CoCounsel over time or expand Thomson's role beyond document review.

AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.