TL;DR
AI companies are facing a tightening data supply as public web text nears saturation and some datasets move behind legal, commercial and national controls. Thorsten Meyer AI’s Control Series says data access is becoming a constraint for AI developers alongside compute and power.
AI companies are entering a phase in which proprietary data, alongside chips and power, is being described as a constraint, according to Thorsten Meyer AI’s Control Series report, which says the public web supply used to train frontier models is nearing its limits while private datasets are being licensed, restricted or kept under national control.
The report cites Epoch AI’s estimate that the public internet contains roughly 300 trillion tokens of high-quality text and says frontier systems are already training on datasets approaching that supply. Epoch’s projection places full use of public human text between 2026 and 2032, with a median around 2028; the report notes that heavier reuse of existing data could move that date earlier.
The report also points to Anthropic’s $1.5 billion settlement with authors as an example of legal and commercial limits on unauthorized scraping. It says the settlement covered past piracy claims involving about 500,000 works at roughly $3,000 per work and required destruction of pirated files, while leaving future training and model-output disputes unresolved.
Thorsten Meyer AI also cites the New York Times’ continuing lawsuit against OpenAI, publisher licensing deals, Meta’s reported $14.3 billion investment for a 49% stake in Scale AI, and Ukraine’s condition that battlefield-related AI arrangements preserve Ukrainian control of the model. The report presents those examples as cases in which some data is behind paywalls, inside companies, held by specialists or controlled by states.
Data: The One Thing You Can’t Rent
The free part of “all human knowledge” is running out. As compute and models commoditize, the corpus you can’t replicate becomes the moat — so data is being fenced, priced, and, in places, treated as a national asset.
Data was supposed to be the abundant input. It’s the scarce one. It’s also the chokepoint you can actually own — so guard your proprietary data, and don’t hand it to a provider who can become your competitor (the lesson everyone fled Scale to learn). Nations: license it like Ukraine — keep the model, keep the leverage.
Private Corpora Become Licensing Focus
The report says compute can be rented and power can be purchased, but exclusive data cannot be copied if another party controls it. It identifies internal documents, customer records, expert annotations, sensor logs and operational histories as data types that may affect model quality in ways that standard cloud infrastructure does not.
For AI startups, paid data regimes can raise the cost of entry. Thorsten Meyer AI says that a settlement or licensing market measured in the billions may favor larger firms with cash, legal teams and publisher relationships, while smaller labs may have fewer ways to train models on protected material.
For businesses outside the AI sector, the report frames data governance as a contractual and competitive decision. Handing proprietary records to a vendor may create short-term product gains, but it can also give that vendor information that helps it build a competing service unless contracts clearly limit training, retention and reuse.

Understanding Open Source and Free Software Licensing
Used Book in Good Condition
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Public Web Text Nears Saturation
The industry grew around the use of large public datasets to train large models. Thorsten Meyer AI says that phase is narrowing because open web text has already been heavily harvested and because rights holders are pushing AI developers toward paid access.
The report says synthetic data is now part of the response. It cites Nvidia’s $320 million purchase of synthetic-data company Gretel and Microsoft’s use of hundreds of billions of synthetic tokens. But the source material treats synthetic data as a partial answer, citing research concerns that repeated training on machine-generated content can compound errors in domains where answers are hard to verify.
That makes verified human data more relevant in legal, medical, scientific and other expert-heavy fields, according to the report. It says the data market has moved away from broad web crawls toward expert-authored, enterprise, real-world and sovereign datasets.
“Public human text is projected to be fully used between 2026 and 2032, with a median around 2028.”
— Epoch AI, cited by Thorsten Meyer AI
private dataset storage solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Legal Boundaries Remain Unsettled
Several parts of the data fight are unresolved. The Anthropic settlement, as described by the report, addressed past piracy claims but did not settle all future training questions, model-output disputes or the wider market price for licensed books.
Epoch AI’s token exhaustion dates are projections, not observed end points. The range could change if model architectures become more efficient, if synthetic data proves more reliable than expected, or if new private and real-world sources become available under license.
It is also unclear how much exclusive data alone will separate leading models if algorithms and inference systems keep improving. The report says data access may become a larger competitive factor as compute and model design become more widely available, but the size of that advantage will vary by field.

AI Data Center Infrastructure Engineering: Power Distribution, Liquid Cooling, High-Density Networking, and Energy Efficiency for GPU Training … Hardware & Compiler Engineering Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Licenses And Data Controls Expand
The next milestones are likely to come through licensing deals, court rulings and enterprise contracts. The New York Times case against OpenAI remains a case to watch, while publisher deals may set prices that smaller AI developers must either pay, avoid or work around.
Companies and governments are also expected to write tighter rules for how their data can be used in AI systems. The report’s practical takeaway is that owners of data that is difficult to replicate will seek contracts that preserve control over training rights, model ownership, retention and downstream use.

AI Compliance & Risk Management for Law Firms: Automated Reviews, Policy Drafting, and Error-Reduction Frameworks: A Comprehensive Guide
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the news in this report?
The development is that AI’s data supply is being treated as a constrained asset in 2026. Thorsten Meyer AI ties that shift to public-web saturation, copyright settlements, licensing deals, enterprise data controls and sovereign claims over real-world data.
Is public internet data actually gone?
No. Epoch AI’s figures, as cited, are projections about high-quality public human text for training frontier systems. The report says the remaining useful supply is being used up, not that the internet has no more text.
Why can’t AI companies just use synthetic data?
They can and do. The report cites synthetic data efforts by Nvidia and Microsoft, but says machine-generated data can create compounding errors when answers are hard to check, which keeps verified human data relevant.
What does this mean for businesses?
Businesses with proprietary records, workflows or expert knowledge may hold data that AI vendors seek to access. The report says giving data to an AI vendor without limits on training and reuse can reduce the owner’s control over that information.
What remains unresolved?
The legal limits on training, the price of licensed datasets, the reliability of synthetic data and the competitive value of exclusive corpora remain open questions.
Source: Thorsten Meyer AI