The Economics of Data Businesses (summary)
Summary of an essay by Abraham Thomas, published in Pivotal on 29 January 2022. Read the full essay on Pivotal →
Abraham Thomas argues that data businesses follow economics unlike any other tech business model: they start slow, then accelerate, and once mature they are almost impossible to displace. Building a proprietary data asset takes heavy upfront investment and a long period of category creation, but unlike software, whose economics tend to degrade, data business economics tend to improve with scale. Thomas, co-founder and Chief Data Officer of Quandl, sets this out as six “fundamental truths”.
What is a data business?
“A company is a data business if, and only if, data is its core product.” Data is central to what the company does; without the data there is no company. Google, Bloomberg, Yelp, and ZoomInfo are data businesses. Every company uses data, but most are not data businesses.
What are the six fundamental truths of data businesses?
| # | Truth | In brief |
|---|---|---|
| 1 | It’s all about the data | Every successful data business is built around a unique or proprietary data asset. |
| 2 | Control unique data to capture unique value | Whoever controls the data captures the value; intermediaries get squeezed. |
| 3 | Data businesses have slow beginnings | Building a usable data asset and its infrastructure takes significant time and money, and usually requires category creation. |
| 4 | Growth accelerates over time | Costs fall, sales get easier, and prices rise as the corpus grows: software economics degrade, data economics improve. |
| 5 | Data businesses are super sticky | Mature data products are almost impossible to displace; churn is minuscule. |
| 6 | Successful data businesses are rare | Valuable data is rare, and the obvious data assets have already been captured. |
How do companies build a proprietary data asset?
Thomas lists eleven methods, which tend to reinforce each other rather than exclude each other:
| Method | Example from the essay |
|---|---|
| Brute force: pay for primary data collection | Google crawling the web; Planet launching hundreds of micro-satellites; ZoomInfo cold-calling company switchboards |
| Aggregate and harmonize others’ published data | Reuters standardizing printed financial statements in the 1970s |
| License and transform commoditized data | Scale AI labelling raw images |
| Affiliate collection by partners with the right incentives | Advertisers installing the Facebook pixel |
| Core business output | Trades on the New York Stock Exchange generating price, volume and order data |
| Payment in kind: a free tool in exchange for data | Foursquare’s free SDKs for app developers |
| Inbound network effects | Site owners submitting data to Google through SEO |
| Give to get | Businesses sending counterparty data to Dun & Bradstreet to access its credit database |
| Data consortia | Like give-to-get, but peer-to-peer rather than hub-and-spoke |
| Data exhaust from the core business | Shopping, email and personal-finance apps that see consumer transactions |
| Data creation (synthetic data) | Tonic generating fake data for testing |
Why do intermediaries get squeezed?
If a business depends on a single upstream source of data, that supplier can raise prices until it captures all the economics. Thomas’s advice is to build a primary data asset, use multiple suppliers, or add enough proprietary value that “a sufficiently large transformation of your source data is tantamount to creating a new data product of your own.” Merging datasets, quality control, labelling, mapping, deduplication, and provenance all add value; dashboards and interfaces alone rarely do. Companies that can neither control data nor add value often pivot to “picks and shovels”: tools that support data businesses.
Why do data businesses start slowly?
- Minimum viable corpus. Almost every data product has a size below which the data isn’t useful. It is the data equivalent of a minimum viable product, but usually much harder to build.
- Parallel infrastructure. Data firms need data ops, data QA, and data product, the equivalents of devops, QA and product in software, and usually have to build them in-house.
- Category creation. Most data products require educating the market, so early sales cycles are long and win rates low.
Thomas treats this slowness as a strength: “my capex is your barrier to entry”, and “a data product that can be built easily is a data product that can be replicated easily.” It also makes data businesses hard to fund, because venture investors index heavily on early growth rates.
Why do data business economics improve with scale?
As a data business grows:
- the marginal cost of acquiring data falls, and infrastructure gains economies of scale;
- sales cycles shorten as the corpus stops being minimal, and a larger corpus reaches more buyers;
- the data can be sliced for targeting and price discrimination, and priced along new axes (per record, per API call, per update);
- the data goes from optional to essential (“the dream of every data asset owner is to become an industry standard”), so prices rise;
- almost all costs are fixed, so more customers spread them further;
- revenue becomes recurring, and net revenue retention and lifetime value become excellent;
- new loops open up: customer contribution loops, data quality loops, data content loops (as used by Expedia, Glassdoor, Zillow, and at Quandl), and data learning loops.
These curves are sigmoid, not unbounded, but they run a long way. The holy grail is to be the lowest-cost acquirer and highest-paying buyer of data inputs while selling data outputs at the lowest price in the market.
What does ZoomInfo show about data business economics?
ZoomInfo began by brute force: its founders called company switchboards to collect contact data, spending about 75% of their time on the phone. A decade later it ran on loops: free access in exchange for users’ email contacts (give-to-get and payment in kind), a customer contribution network, a machine-learning data quality loop, customer exhaust data such as email bounces, and SEO-driven data content. The results, as reported in the essay: sales cycles under 30 days, margins rising from 50% in 2014 to 90%, net revenue retention above 100%, a 6–8 month payback period, and a 15:1 LTV-to-CAC ratio. ZoomInfo was bootstrapped until a private equity round in 2014.
Why are mature data businesses so hard to displace?
- Disruption from below rarely works: a cheaper product still has to reach a minimum viable corpus, and data acquisition economics make that hard.
- “Better” is hard to define for data: a challenger usually needs to be an order of magnitude faster, more accurate, or more comprehensive, and even that may not raise the data’s usefulness enough.
- New winners use new data: successful new data products usually come from a different source and a different set of loops, while the original asset keeps selling.
- Incumbent tactics: commoditizing the complement, creating data standards (from PCI to CUSIP), and buying adjacent data.
The biggest data businesses tend to form duopolies: Visa and Mastercard, Nasdaq and NYSE, Moody’s and S&P, Experian and Equifax, Bloomberg and Refinitiv. They also last. Dun & Bradstreet is about 180 years old, and four US presidents (Lincoln, Grant, Cleveland, and McKinley) worked for it. Equifax dates from 1899, Reuters from 1851, and Standard & Poor’s from 1860.
Why are successful data businesses rare?
Valuable data is rare (Sturgeon’s law: most of everything is junk), and obviously valuable data about companies and people has already been captured. That leaves non-obvious data assets, which require category creation. Thomas speculated that data might see a “deployment era” of targeted, vertical, application-specific data businesses rather than new horizontal platforms.
Related
- Full essay: The Economics of Data Businesses, Pivotal, 29 January 2022
- How to Price a Data Asset (summary)
- Data and Defensibility (summary)
- Quandl: founders, alternative data, and the Nasdaq acquisition