Public Data for AI Is Running Out. Your Company's Own Data Just Became the Moat.
Blog

Public Data for AI Is Running Out. Your Company's Own Data Just Became the Moat.

Public Data for AI Is Running Out. Your Company's Own Data Just Became the Moat.

There's an assumption baked into most conversations about AI that nobody bothers to state out loud: that the raw material feeding these models, the internet's endless supply of written text, is effectively infinite. It isn't, and the timeline for that becoming a real constraint is a lot shorter than most people realize.


Researchers at Epoch AI, presenting peer-reviewed work at the 2024 International Conference on Machine Learning, estimated the entire effective stock of high-quality, human-generated public text available for training at roughly 300 trillion tokens, with a wide confidence interval running from 100 trillion to 1,000 trillion. Their original projection put the exhaustion point for the highest-quality portion of that pool as early as this year. Since then, the estimate has shifted slightly later, into 2028 and beyond, as researchers accounted for a handful of untapped high-quality sources like digitized library archives. But the direction of the finding hasn't changed. The free, open internet's supply of genuinely useful, high-quality text is a finite resource, and the industry is closing in on the bottom of it.


Why More Compute Doesn't Solve This

The instinct is to assume this is a problem that scale eventually fixes. It's the opposite. The practice of "overtraining," running a model on far more data relative to its size than the classic compute-optimal formula calls for, has become standard because it produces better results per dollar of inference cost. Meta's Llama 3 8B model, for instance, was trained on roughly 100 times more data than the compute-optimal amount for a model of its size. That approach produces better models. It also burns through the finite pool of quality data dramatically faster than raw model count alone would suggest.


This is why the AI labs most exposed to the shortage aren't waiting around. OpenAI has struck data licensing deals with Reddit and News Corp, among others, explicitly to secure a flow of material that isn't available for free anymore. That's a meaningful signal in itself. When a company with access to virtually unlimited engineering talent and capital starts paying for the raw material it used to scrape for free, the scarcity is real, not theoretical.


The Part That Actually Matters for Everyone Else

None of this is really a story about AI labs, though. It's a story about what happens to the value of data everywhere else once the free public pool stops being the differentiator it once was.


Every business generates data that never touches the public internet. Customer interactions. Internal process documentation. Historical transaction records. Support tickets. Proprietary workflows refined over years of operating in a specific market. None of that has been scraped, indexed, or absorbed into a foundation model, because it was never public to begin with. As the industry's access to fresh, high-quality public text tightens, the relative value of that private, ungenericized operational data climbs, simply because it's one of the few remaining sources of information a model hasn't already seen a thousand variations of.


This reframes a question a lot of leadership teams have been asking backwards. The question hasn't been "how do we get access to more AI." Every competitor has access to roughly the same foundation models. The real question is what a business can bring to those models that nobody else can, and increasingly, the honest answer is: whatever data only that business actually has.


Why Most Companies Can't Actually Use the Advantage They're Sitting On

Here's the complication. Having proprietary data and having usable proprietary data are two very different things. Most businesses' internal data lives scattered across disconnected systems, inconsistently structured, only partially governed, and rarely captured with the kind of consistency that makes it trainable, retrievable, or even reliably queryable.


Research into what's actually reshaping data engineering in 2026 keeps landing on the same reframe: the strategic focus is shifting away from raw data volume and toward data quality and, especially, data freshness. A model or retrieval system built on top of stale, poorly structured internal data doesn't get to skip the same failure modes plaguing public-data-trained models just because the data is proprietary. Bad data is bad data regardless of where it came from. The advantage only materializes once that data is captured cleanly, structured consistently, and kept current enough to actually be useful.


That's the gap between businesses that will benefit from this shift and businesses that will watch it happen to someone else. It isn't about who has more data. Most established businesses, especially ones that have operated for any real length of time, have plenty. It's about whether that data has ever been treated as something worth investing engineering effort into, rather than a byproduct that accumulates in whatever system happened to generate it.


What This Actually Changes

The businesses paying close attention to the AI data shortage right now tend to be the ones building foundation models, and that makes sense, it's their direct problem. But the more interesting long-term shift is happening one level removed from that conversation.


As the public data pool that every foundation model was trained on becomes increasingly saturated and repetitive, the differentiation between what one company's AI-powered workflows can do and what a competitor's can do increasingly comes down to what gets fed into those workflows on top of the shared foundation. A generic model with access to a decade of clean, structured, genuinely proprietary operational data will consistently outperform the same generic model working with scattered spreadsheets and half-documented processes, on the exact same underlying architecture.


The scarcity story in AI right now is framed as a problem for the handful of companies building frontier models. For nearly everyone else, it's closer to an opportunity that most businesses aren't yet positioned to take advantage of, because the data that would matter most has never been treated as an asset worth structuring properly. That gap won't stay open indefinitely. The businesses that close it early are the ones building on genuine ground, not competing over an increasingly picked-over public well alongside everyone else.

Share this blog post