Perpetual Data Machines
I’m seeing a lot of talk on the timeline about AI data brokers - their incredible revenue ramps, but also concerns that their cashflows are temporary, and therefore what’s the terminal value, and how do you finance them. I want to clear up one popular misconception about these businesses:
A particular dataset may have zero terminal value. This does not imply that the accompanying data business has zero terminal value.
There are at least four ways to build perpetual data machines.
Perishable Data
If your dataset has an expiry date, you can build a good business just selling updates.
Most “classic” data-as-a-service businesses follow this template. The revenue is genuinely recurring, but has industry-specific nuances - financial data has a short half-life, and hence is very lucrative; adtech and profile data, less so.
AI training data doesn’t fit this pattern. It’s not very perishable; once you’ve trained, you’ve trained. The marginal value of an update to an existing data corpus is usually minimal. So this is usually NOT how you build a sustainable AI data franchise. (Most hesitations about the economics of AI data brokers stem from people applying DaaS thinking to the category.)
Hunters for Hire
Build a relationship with a lab. Learn how to deliver data to their spec. Become a trusted vendor. Understand their model roadmap. And then go out and build the dataset that feeds it. Effectively, you become an outsourced data team for the lab (or ideally, labs).
But this is not enough. Anybody can be a hunter for hire - there’s no moat or sustainable business there. The trick is to build the infrastructure (tech, people, networks) to do this repeatably, at scale, at the required quality, and above all fast.
This is what Scale, Mercor, Handshake et al do so well: they skate to where the puck is going, and they stand up new data products incredibly quickly. There’s also a new generation of synthetic data providers that do the same thing, but with machines instead of people.
With these businesses, although the revenue appears to be for the data, it’s actually for the infra that underlies the data. And hence it’s stickier than you might think.
Magnets in the Middle
The classic marketplace play: connect buyers and sellers of data.
Marketplaces are notoriously not recurring revenue: if your product isn’t perishable (subscription), you have to reacquire customers and vendors from zero every month. Going from zero to one is hard.
But at-scale marketplaces are terrifically long-lived businesses. Their sustainability comes from magnetism: the buyers go where the sellers are, and the sellers go where the buyers are. If you can consolidate fractured supply, unlock brand new (often untapped) sources, or layer in a superior delivery experience, you end up with multiple network effects that are very hard to disrupt.
The marketplace business model works even if any particular dataset has zero terminal value. But it’s hard to do well.
(I should know; I co-founded Quandl, a data marketplace that was acquired by Nasdaq. Quandl had perishable financial data; we were trusted data hunters for all the blue chip quant funds; and we had marketplace dynamics on our side – but it was still an insanely challenging business to build.)
Guard your Gold
Flip the equation. The labs, who have more information than you, think that data is so valuable that they’re willing to pay sums of money that make AI data brokers the fastest growing companies in history (by revenue). This is a strong indicator that you should not sell data to the labs; you should capture its value yourself.
This is of course what Meta is doing with Muse (ooh, I wonder who they hired to run that business). But even at smaller scale, this is the smart and ambitious play.
And this too is a perpetual data machine, iff you can close the loop. Better data leads to a better product; better product leads to more usage; more usage leads to better data. SaaS data learning loops tend to plateau, but the lesson of recent years is that AI learning loops do not!