Skip to content
← Back to Insights

Thomson Reuters Builds Its Own Legal AI Model: What $40 Million Buys

AI LLM Legal Tech Enterprise AI News

TL;DR

Thomson Reuters has launched Thomson, its proprietary large language model, beginning with tabular analysis in CoCounsel Legal. The company spent $40 million, while a full technical report and broad independent validation remain outstanding.

Thomson Reuters Builds Its Own Legal AI Model: What $40 Million Buys

One measurement can test this launch over the next three to six months. If tabular analysis in CoCounsel Legal reduces inference expense and human correction time on the same document sets and review standards, Thomson Reuters will have changed the economics of its product; the company has not yet published that evidence.

On 2026-08-24, Thomson Reuters launched Thomson, its first proprietary large language model, and disclosed a combined talent and compute investment of $40 million. It did not train a foundation model from scratch. The company started with an open-weight model, then applied mid-training and post-training using material and tools from Westlaw, Practical Law, Checkpoint, and Reuters. Hundreds of domain specialists helped define objectives, construct examples of legal questions, and judge answers in blind comparisons. The expenditure therefore bought specialization and control rather than a general-purpose race for the largest possible model.

The first deployment is Tabular Analysis in CoCounsel Legal, a feature for high-volume, structured document review. Administrators will still be able to select other models, and CoCounsel remains a multi-model system. Thomson will handle work where the company sees a domain advantage, while outside frontier models will continue to serve other tasks. That design limits the risk of moving every workload at once and creates a practical setting in which Thomson Reuters can compare the quality and cost of its own model with those of suppliers.

Cost is central to the decision. SiliconANGLE reports that the project ran for two years and that the final training run cost about $450,000. The company also expects ownership to reduce inference expenses. Less than 10% of the firm’s information base has been used so far, and its researchers say the next step is not simply adding documents. They intend to turn useful content and product activity into better training signals. Thomson Reuters also says customer information will not be used for training without explicit consent.

Internal results still need reproduction

According to the company, internal tests found Thomson broadly competitive with leading models when all systems had access only to the public web. When connected to Thomson Reuters content, Thomson moved to roughly equal or slightly better results. The evaluations considered both answer completeness and whether citations supported the claims they accompanied. However, the full technical report has not been published, and SiliconANGLE notes that the findings have not received extensive independent validation. Early assessments by two legal academics add outside observations, but their sample size and conditions cannot establish performance across jurisdictions and types of legal work.

The company plans to give more external parties access over the coming weeks and months. It also intends to release a smaller open-weight version on Hugging Face for academic and non-commercial use, while extending Thomson models into legal and tax products and adding sovereign-AI options. By the end of 2026, the most informative evidence will not be a single benchmark announced by the vendor. It will be the human correction rate in Tabular Analysis, the cost per reviewed document, and whether outside researchers can reproduce citation quality under disclosed conditions.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.