This is going to increasingly happen over the years to come. Big organizations will become more sophisticated with operationalizing their data, training and running LLMs will continue to be demystified and accessible, and over time we'll get more and more specialized / industry-specific models.
It's going to become another way to monetize your informational assets if you're a big older enterprise with troves of data. All you need is time to figure out how to make it useful for yourself and then eventually sell access to it however you want.
Think of all the data that big orgs have that isn't accessible to all the AI labs to suck up.
> Our evaluation found Thomson’s citation quality generally competitive with leading frontier models, even when tested on Canadian employment-law questions without a Canada-specific setting.
That’s it? It was generally competitive with leading frontier models? Neat, but why would someone pay for frontier models and also a generally competitive additional product?
Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models.
It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.
> Thomson Reuters is also making a “small” version of Thomson available as an open-weight model on Hugging Face for academic and non-commercial use to further aid in this validation.
I don't trust that they'll be able to make back that $40M.
This feels very much like a news agency getting into crypto or launching its own NFT line.
Or IBM selling Watson.
Or Mozilla chasing every which thing.
They're not stakeholders in the future of work. They're just wanting to stay relevant and pattern matching against what they see.
Reuters is too important for this.
If they were trying to use this as a narrative affront to OpenAI and Anthropic, maybe, but this is Reuters, not a deeply political organization seeking to land gotchas against big tech.
Keep in mind line items can be deceptive. e.g. “Corporates” can mean “money we make from selling access to individual’s biometric data to the government.”
1. Marketing and expressing to their customers that they are not falling behind, and
2. Insulating themselves from frontier labs jacking up prices, nerfing the models they depend on, or otherwise unexpected changes in behavior.
I think the main goal is #2. Thomson Reuters might be a $40B company, but.... at this point it's not clear that that holds any weight in terms of not being fucked over by 2 companies aiming for $2t+ IPO valuations.
Edit: On second thought, there is probably a #3 too. They can serve inference for their own models significantly cheaper than frontier lab rates (assuming they're capturing continuous use of their hardware). I still think #2 is the primary goal.
No it's just a value add to their existing data products and a moat against the big n LLM companies. The "news" part of the business is relatively small compared to everything else the do.
Like how Bloomberg does news but its far from their only or primary product. TR covers a different surface of data products than Bloomberg but it's a decent comparison.
Here is some more technical information on how this was trained, as well as a download link.
https://huggingface.co/thomsonreuters/Thomson-1.0-Small
(Full disclosure I’m a TR employee, although I had nothing to do with making this)
This is going to increasingly happen over the years to come. Big organizations will become more sophisticated with operationalizing their data, training and running LLMs will continue to be demystified and accessible, and over time we'll get more and more specialized / industry-specific models.
It's going to become another way to monetize your informational assets if you're a big older enterprise with troves of data. All you need is time to figure out how to make it useful for yourself and then eventually sell access to it however you want.
Think of all the data that big orgs have that isn't accessible to all the AI labs to suck up.
> starting from a strong open-source foundation and investing $40 million to train Thomson
Sounds like they spent $40 million finetuning an open weight model on their own data? I wonder what they built on.
Per Business Insider [1], it was built on Qwen.
[1] https://www.businessinsider.com/thomson-reuters-builds-ai-mo...
> Our evaluation found Thomson’s citation quality generally competitive with leading frontier models, even when tested on Canadian employment-law questions without a Canada-specific setting.
That’s it? It was generally competitive with leading frontier models? Neat, but why would someone pay for frontier models and also a generally competitive additional product?
It’s a cost-saving measure and a marketing story.
Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models.
It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost inefficient compared to using an API. Idk how Reuters will manage to solve this fundamental problem.
> Thomson Reuters is also making a “small” version of Thomson available as an open-weight model on Hugging Face for academic and non-commercial use to further aid in this validation.
Looking forward to the ERP fine-tune.
Has anyone found a link to the technical report? They don’t seem very keen to publicise their performance on evals…
It’s on hugging face, I suppose one could do those themselves?
https://huggingface.co/thomsonreuters/Thomson-1.0-Small
I don't trust that they'll be able to make back that $40M.
This feels very much like a news agency getting into crypto or launching its own NFT line.
Or IBM selling Watson.
Or Mozilla chasing every which thing.
They're not stakeholders in the future of work. They're just wanting to stay relevant and pattern matching against what they see.
Reuters is too important for this.
If they were trying to use this as a narrative affront to OpenAI and Anthropic, maybe, but this is Reuters, not a deeply political organization seeking to land gotchas against big tech.
Reuters, the news agency, is only roughly 10% of TR's revenue. They're still a pretty big player in Legal and Tax.
Keep in mind line items can be deceptive. e.g. “Corporates” can mean “money we make from selling access to individual’s biometric data to the government.”
https://ir.thomsonreuters.com/news-releases/news-release-det...
I doubt the point is to "make the money back".
It's likely split between two goals:
1. Marketing and expressing to their customers that they are not falling behind, and
2. Insulating themselves from frontier labs jacking up prices, nerfing the models they depend on, or otherwise unexpected changes in behavior.
I think the main goal is #2. Thomson Reuters might be a $40B company, but.... at this point it's not clear that that holds any weight in terms of not being fucked over by 2 companies aiming for $2t+ IPO valuations.
Edit: On second thought, there is probably a #3 too. They can serve inference for their own models significantly cheaper than frontier lab rates (assuming they're capturing continuous use of their hardware). I still think #2 is the primary goal.
No it's just a value add to their existing data products and a moat against the big n LLM companies. The "news" part of the business is relatively small compared to everything else the do.
Like how Bloomberg does news but its far from their only or primary product. TR covers a different surface of data products than Bloomberg but it's a decent comparison.