Eric Schmidt, who was the CEO of Google from 2001 to 2011, and an early investor in Anthropic, gave a talk to Stanford in April 2024. The talk became controversial and was taken down.

One of the insightful remarks that he made was about how LLMs have already been trained on all the public data, and he wondered how these LLMs will get more of it to improve.

The answer became obvious in 2025 when Facebook/Meta acquired Scale AI (via HALO).

Scale AI provides contractors who interacted with LLMs in specialized areas and created new synthetic test data for training. It makes $1 Billion+ a year which says something about the hunger that LLM providers have for data.

There are several such companies in this space now.

One thing I recently learnt is that this training data, per se, is not proprietary to their customer. So, if a contractor finds a useful evaluation case while evaluating Anthropic, the company can resell the same evaluation to Google Gemini or OpenAI as well.

And that’s how all the new specialized data which is used for training and evaluating LLMs gets shared across major LLM providers in the US.

In a way, it is a form of indirect distillation.