How to Price a Data Asset (summary)
Summary of an essay by Abraham Thomas, published in Pivotal on 11 May 2024. Read the full essay on Pivotal →
How do you price a data asset? Abraham Thomas’s answer is that data has no innate value: “the value of data is the value of the marginal change in actions taken after adding the data to your business process.” Price therefore depends on the use case, the user, the dataset’s position in its lifecycle, its uniqueness, its quality, and the usage rights granted. Standard software pricing axes such as seats, features and tiers mostly fail for data.
Thomas bases this on his time as co-founder and Chief Data Officer of Quandl, where he evaluated thousands of data assets and priced hundreds of data products. He writes that he has priced “more — and more varied — data products than almost anyone in the world.”
Why is data hard to price?
Data is inherently heterogeneous. Crude oil has agreed transaction criteria (volume, location, grade, date) and benchmark contracts; data has none. “Every barrel of WTI crude is identical; no two datasets are identical.” There is no single formula, but there are principles that generalize.
Historically, two industries dominated data purchases: finance and adtech, the only ones with many buyers, many valuable datasets, and the ability to pay material recurring revenue. The essay argues there is now a third buyer, AI, whose needs and value curves differ, so much past pricing intuition no longer applies.
What are the axioms of data value?
| Axiom | Meaning |
|---|---|
| Data has no innate value | Its value is the value of what can be done with it. |
| Value depends on the use case | Financial statements are useless for an ad campaign, and audience profiles for equity analysis; flip them and each is essential. |
| Value depends on the user | A scale effect (more capital, customers or parameters) and a capability effect (sophisticated users extract more). Much of pricing is finding proxies for both. |
| Data is additive | Unlike software, where a second CRM adds nothing, more records, fields or overlapping sources make data more valuable. |
| Data is rivalrous … | Contrary to the common view, only one hedge fund can profit from unpriced information in a dataset; owners of valuable data keep it exclusive. |
| … until it isn’t | Advantages decay, substitutes appear, and data becomes non-rival. |
| Data has a lifecycle | Where a dataset sits in its lifecycle drives its price (see below). |
| Marginal lift is what matters | Value is the marginal change in actions the data causes. |
What is the lifecycle of a data asset?
| Stage | What happens | Price |
|---|---|---|
| 1. Early | The dataset and the market are both immature; transactions are rare | Low; exploratory |
| 2. Early adopters find “alpha” | A few users get an edge from it | Very high, narrow audience |
| 3. Widespread | Substitutes proliferate and the edge decays | Declining |
| 4. Table stakes | Not using the data puts you at a disadvantage | Rises again, with much broader usage |
In capital markets, Thomas places unstructured training data between stages 1 and 2, much alternative data in stage 3, and market data in stage 4. Table stakes is the best position for a data owner: “The dream of every data owner is for their data to become table stakes — commoditized, but essential.”
Why is unique data so valuable?
Unique data is universally additive; its owner can control the pace of its lifecycle to maximize price times transactions; and if it becomes table stakes, the owner effectively collects a tax on an industry. The catch is functional substitution: very different datasets can deliver the same insight. Foot traffic, email receipts and credit card transactions all reveal what people buy at the mall. Among substitutes, the most valuable is the one “closest to the sun”, most tightly linked to the event of interest.
Which pricing axes work for data?
| Pricing axis | Works for data? | Why |
|---|---|---|
| Per seat | No | Data value doesn’t scale with user count; data is used by teams. |
| Per feature | No | There is rarely a meaningful “data feature” slider. |
| Raw volume (terabytes) | Usually no | Only for fungible, commoditized data. It now works for some AI training data. |
| Per API call | Usually no | Only if data decays fast, or the call triggers an action. |
| Per download | No | Data is trivial to copy, and auditing is hard. |
| Structured volume | Yes | More records, fields, coverage, history or granularity. |
| Quality | Partly | Quality has several dimensions. |
| Access | Yes | Speed, recency, update cadence, exclusivity, usage rights. |
| Use case | Yes | Possible for data, unlike software. |
| Customer scale | Yes | Larger customers get more value. |
| Business unit | Yes | Team, geography, product line or model generation. |
Two ways around the problem are to wrap the data in software, selling the most valuable use case as an application (Google’s ad business on intent data, Experian on credit data, the Bloomberg Terminal, ChatGPT on model weights), or to wrap a service in data (Scale AI, Clearbit, Datavant): “perform the service once, but sell it many times.”
How does data quality affect price?
Quality depends on the buyer:
- Hedge funds care about accuracy and precision, because bad points look like false signals.
- Adtech cares about coverage and depth.
- AI model builders care about structure: deduplicated, denoised, debiased data, annotations, diversity, and some domain specificity.
The same dataset can therefore be priced differently for each, and many quality attributes can be improved to raise value. Other internal drivers of value are provable compliance, provenance (primary sources are the gold standard), uncontaminated data (“a non-renewable asset”, since any use for training or testing spends it), and fungibility.
How does data become table stakes?
- Exchange standards: CUSIP, DUNS, Datavant’s patient key, LiveRamp’s RampID.
- Evaluation benchmarks: market indices from S&P, Nasdaq and MSCI; Nielsen ratings.
- Quasi-monopoly: Meta’s and Google’s knowledge graphs; Bloomberg’s terminal.
- Bundled usage: realtors and the MLS databases.
Usage rights are priced separately: scope of use, ownership, audit rights, derived-data rights, and compliance protections all carry dollar values.
How does AI change data pricing?
- Quantity matters, a lot. For AI there seems to be no upper limit to the value of more data, so paying by volume now makes sense, and Sturgeon’s law (“90% of everything is junk”) weakens. Beyond a point, though, “data quality scales better than data quantity”: cheap improvements such as deduplication beat acquiring more tokens.
- Recurring revenue is hard. Most training value lies in the historical corpus, not ongoing updates. News media is the exception, with deals split into a fixed archive fee and a variable fresh-content fee. The long-term answer is a data flywheel that keeps generating new data.
- Synthetic data lowers the value of existing data (“proprietary data ain’t what it used to be”), though synthetic data can degrade over time, which gives provably human data a lasting place.
- Payment in kind is emerging: AI companies pay publishers partly in brand placement and traffic.
What else determines data prices?
- Sales cycle determines ACV. Buyers spend expensive engineering time testing a dataset, and that effort signals value, so sellers raise prices. It is also why messy datasets are often more expensive.
- Legibility determines market size. The easier a dataset’s ROI is to compute, the larger its market; that is why finance and adtech are the most lucrative data verticals.
Related
- Full essay: How to Price a Data Asset, Pivotal, 11 May 2024
- The Economics of Data Businesses (summary)
- Data and Defensibility (summary)
- On Data Quality (summary)