Thomson Reuters builds its own legal AI model instead of renting Anthropic or OpenAI's
The $40 million system is fine-tuned on the company's own legal and news data, reflecting a broader shift toward owning AI infrastructure rather than paying frontier labs by the token.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
- Thomson Reuters has built its own AI model, called Thomson, fine-tuned on the company's proprietary legal, tax and news content rather than built entirely on a rented frontier API, The Decoder reported on August 24.
- The company has invested $40 million in talent and compute on the project and has so far trained the model on less than 10% of its available content, according to the report.
- The Decoder reports the model is built on an open-source foundation it identifies as Qwen3.5-397B, a detail Thomson Reuters has not publicly confirmed.
Thomson Reuters has built its own AI model, called Thomson, trained on the company's proprietary legal, tax and news content — a move the company frames as owning its AI infrastructure rather than renting capability from Anthropic, OpenAI or other frontier labs, The Decoder reported on August 24. The outlet reports the company has invested $40 million in talent and compute on the project.
The Decoder reports Thomson Reuters has trained the model on less than 10% of the proprietary content it has available, implying room to keep specializing it further. Thomson Reuters has not named the open-source model Thomson was built on; The Decoder reports, citing sourcing this outlet did not independently verify, that the base is Alibaba's Qwen3.5-397B.
The Decoder also reports specific benchmark comparisons — including a Stanford LegalBench score of 0.823 that trails Gemini 3.1 Pro and GPT-5.5, and a Harvey Legal Agent Benchmark result described as "just behind" Anthropic's Opus 4.8 — that should be read as one outlet's reporting rather than a company-confirmed figure.
A company the size of Thomson Reuters concluding it is cheaper and more defensible to fine-tune an open model on its own data than to keep paying frontier labs by the API call is a real data point on the build-versus-rent calculation every enterprise with a large proprietary corpus is currently running — and if the underlying model choice really is a fully open one, it also means the AI labs' moat looks thinner to a well-resourced customer than their pricing implies.
Will Thomson Reuters confirm or deny the open-source base model The Decoder identified as underlying its new Thomson system?
Still open. When the paper finds out, it will say so here and on the open questions page — including if it got this wrong.