Data in the Age of AI (summary)

Last updated: 29 September 2026

Summary of an essay by Abraham Thomas, published in Pivotal on 20 May 2023. Read the full essay on Pivotal →

How does AI affect data and data businesses? Abraham Thomas’s 2023 answer has two parts. First, data and software are complementary inputs, so when AI makes software (“compute”) dramatically cheaper and more productive, data becomes relatively scarcer and more valuable. Second, generative models don’t just consume data, they produce it, so the total amount of content explodes, and trust, provenance, identity, quality, and curation become essential. He calls the chain of proofs needed to trust any piece of data “the confidence chain”.

What was the data explosion?

Falling storage costs, and services such as Amazon S3 that made storage easy to use, made the 2010s “the decade of the data explosion”. At Quandl, Thomas notes, “we saved everything.” The decade’s winning business models were built on cheap, plentiful data: the content, adtech and social ecosystem (users, content and advertisers in one loop) and the e-commerce, delivery and logistics ecosystem (Amazon and Uber flywheels are data learning loops). These flywheels generate more data in turn: “data begets data.”

Why does cheaper software make data more valuable?

“Data and software are two sides of the same coin”: software is useless without data, and data is worthless without software to act on it. Economists call these perfectly complementary inputs. When one gets much cheaper, the price of the other tends to rise. Thomas’s analogy: if needles become cheap, tailors buy more fabric, and fabric prices rise.

For a decade, data was abundant and software was the scarce input, which is why engineers and software companies commanded high prices. LLMs (the “compute explosion”) flip this, since every programmer can be far more productive. The first consequence: “data just got a whole lot more valuable.”

Thomas calls LLMs “the peace dividend of the content wars”: the transformer architecture was invented at Google to cope with the flood of web content, and like Cold War technologies it will be used far beyond its original purpose.

Which kinds of data gain value with AI?

Kind of data Example from the essay
Unique data assets BloombergGPT, trained on decades of proprietary financial data (“a twenty year lease of life” for Bloomberg, per an anonymous industry executive)
Latent data assets, valuable but previously unmonetized Reddit’s human-generated content, now behind a paid API
Small custom data Fine-tuning techniques such as LoRA let modest proprietary datasets add value to large models
Golden data, of exceptional quality for a use case “Data quality scales better than data size” above a certain corpus size

There are also “picks and shovels” opportunities: tools to build new data assets for AI, connect existing assets to AI infrastructure, extract latent data using AI, and monetize data assets. More broadly, the data stack must be rebuilt so that “generative models become first-class consumers as well as producers of data”, with new pricing models, data rights, compliance, and data marketplaces. “No more ‘content without consent’.”

What is the confidence chain?

In a world of unlimited content, legitimate and otherwise, trust depends on a chain of proofs “only as strong as its weakest link”: signatures, provenance, identity, quality, curation. Who created this, can you prove it, are they who they say they are, is it good, and does it match what I need? Signatures, provenance and identity remain technologically open (Thomas notes it is an ideal use case for zero-knowledge cryptography). Quality and curation are being handled by emerging trust hierarchies. His 2023 guess at the order:

filter bubbles > friends > domain experts ≈ influencers > second-degree connections > institutions > anonymous experts ≈ AI ≈ random strangers > obvious trolls

He expected “curated AI” might move up. Curation, he argues, is about matching, not ranking. If AI multiplies content a hundredfold, quality filters don’t need to be a hundred times stricter; the limit is the consumer’s bandwidth, so the goal is the best match above a minimum quality bar.

What does “compute all the things” mean?

Just as the default for data flipped from “conserve memory” to “save everything”, the default for software will flip to “compute everything”: “agents, agents everywhere.” Instead of humans in the loop improving software, “software-in-the-loop” will streamline human processes, through co-pilots, research assistants, tutors, and personal curators.

Where will the new scarcities be?

Thomas applies the Jevons paradox: as “informed computation” gets cheaper, society uses much more of it. That makes chips scarce, from insatiable demand rather than insufficient supply (the essay cites NVIDIA’s rise). It also makes energy a likely constraint, and he warns of artificial scarcity created by firms trying to capture the gains. He later noted that this essay “may also have been one of the earliest to cite Jevons’ Law while analyzing NVDA.”