Data and Defensibility (summary)
Summary of an essay by Abraham Thomas, published in Pivotal on 12 April 2025. Read the full essay on Pivotal →
When does data create a moat? Abraham Thomas argues there are exactly two categories of data moat, data control and data loops, and that much popular thinking about them is wrong. Unique data is “neither necessary nor sufficient” for a moat, and data learning loops, the most familiar kind, are usually not moats at all. The strongest are systems of record and action, industry standards and benchmarks, clearinghouses for fragmented data, catalyst data, and data gravity. AI weakens some of these (brute-force data collection, SEO loops, system-of-record lock-in) and strengthens others.
Why do data moats matter more with AI?
AI startups are growing faster than ever: the essay notes Bolt reaching $20M ARR in two months, and Cursor going from $1M to $100M ARR in 21 months. But competition is brutal, and new foundation models can erase capabilities overnight. Data and AI are “two sides of the same coin”: models need data, and models unlock data’s value. So “data moats reinforce AI advantages, and AI advantages reinforce data moats.”
What are the two kinds of data moat?
- Data control: sole control over the production, movement, or usage of valuable data.
- Data loops: positive feedback cycles in which data improves the business and the business improves the data, fast enough that competitors can’t catch up.
Every data advantage fits one or both.
When does unique data create a moat?
Only when it is meaningful, which requires all three of:
- substantial value to you or your customers;
- genuine rivalry, so that others can’t get the same value from it;
- no functional substitutes, meaning no other data can achieve similar outcomes.
Most datasets meet none of these. Exhaust data, such as NYSE’s and Nasdaq’s market data, is not a moat for the core business. Process power (FactSet, Moody’s, Nielsen) is a moat. Brute-force collection was one, but is weakening.
Is brute-force data collection still a moat?
Mostly not, for three reasons: LLMs make data acquisition “orders of magnitude easier”; capital for data acquisition is cheap; and knowledge and tools diffuse. Upstarts can now “replicate 99% of their work for 1% of the cost.” Brute force still works where the last 1% of accuracy or coverage matters (for example, finance), upstream of LLMs (labelling, synthetic data, evals), and in real-world domains such as audio, video, physics and biology, though “this won’t last”. A funding advantage can still buy a first-mover lead, to be converted into another moat.
Which kinds of data control create moats?
| Type of data control | Examples in the essay | Verdict in the essay |
|---|---|---|
| Clearinghouse for fragmented data | Bloomberg, LexisNexis, CoStar; also Rippling, Stripe, Plaid, Zapier, Kayak | A genuine moat; best when the data is a substrate for action and needs a one-to-many relationship with the use case |
| Controlling external data movement | Visa, Amadeus and Sabre, Change Healthcare | Lucrative but rare |
| Managing internal data movement | Most software; exceptions in regulated industries (Epic, Datavant) | Usually no moat, except under regulation or complexity |
| System of record (SoR) | Salesforce, Oracle, Workday, SAP, Epic, QuickBooks | One of the oldest and most effective moats; now threatened by AI-driven migration |
| System of action (SoA) | GitHub and GitLab | Stronger than an SoR, because the action layer coheres with the data and the user’s function |
| Systems of agents | Coding agents such as Cursor, Devin, GitHub Copilot | Emerging: the system acts on the data itself |
| Exogenous control: IP, contracts, regulation | S&P 500 licensing (about $1B a year); IQVIA; ENERGY STAR certifiers | Real; contractual monopolies are temporary, government-backed ones inertial |
| Catalyst data | Google intent data, Amazon purchase history, CUSIP, DUNS, FICO, Nielsen | Tends toward winner-take-most |
Why systems of record are sticky: they control data usage, so nothing can be done without them, and years of “workflow barnacles” accumulate around them. The founding years show how durable they are: SAP 1972, Oracle 1977, Epic 1979, QuickBooks 1983, Salesforce 1999. But exporting data from an SoR is exactly the kind of tedious task AI agents excel at. Disruptors now offer to do the migration themselves, and “data viscosity is your friend, until it isn’t.”
Systems of record, action, and agents: “Systems-of-record store data; systems-of-action empower humans to act on data; systems-of-agents act on data themselves.”
Catalyst data is data that “activates”, that is materially increases the value of, other data: Google’s user-intent data activates search results, and unique identifiers activate siloed records.
Which data loops are moats?
| Loop | How it works | Verdict in the essay |
|---|---|---|
| User-generated content | Content attracts users, users attract advertisers (Facebook, YouTube, TikTok) | Deceptive: it “can reverse just as fast as it builds” (MySpace, Tumblr, Vine) |
| SEO loop | Useful content wins search traffic, which brings more data (Zillow, Yelp, Glassdoor) | Strong from about 2005 to 2020; now eroding through AI slop, LLMs replacing search, and paywalls. Stack Overflow’s traffic peaked in May 2020, just before GPT-3 |
| SaaS data gravity | The tool holding the most central data swallows the others (Toast) | “One of the best data moats out there” in vertical SaaS |
| Give-to-get | Contribute data to access the aggregate (Waze, Verisk, Dun & Bradstreet, PayScale) | Moaty once past critical mass; hard to cold-start |
| Data learning loop | Using data to run better, which yields more data | “Not a moat”: value plateaus while costs rise |
| Secondary learning loops | Data quality, recommendations, A/B testing | Not moats (“A/B testing is not even a shallow ditch”) |
| Usage/value loops | Exchange standards (CUSIP, DUNS, VIN, ISBN), evaluation benchmarks (S&P 500, Nielsen, FICO), pass-through (Stripe fraud data), trust (Moody’s, Verisign, Elsevier), implicit knowledge capture | Possibly Thomas’s favourite; a data asset that grows more valuable the more it is used “mints money” |
Exceptions for learning loops. A learning loop can be a moat when it unlocks a business model that is otherwise impossible (Amazon Prime’s free two-day delivery), inside data businesses, and, partly, in AI, where outcomes scale with data but there is no ongoing loop yet. Otherwise learning is best used in a bootstrap and switch: grow with a learning loop, then build network effects or other moats at scale (Facebook, Netflix, Uber, Airbnb).
Implicit knowledge capture is an emerging usage/value loop for AI application companies. Start from a foundation model, automate a vertical’s workflows, learn its edge cases from human feedback, earn trust, widen the scope, and “become irreplaceable”. The limit is the complexity of the real world.
Are data network effects real?
True data network effects “are rarer than once thought.” Data scale effects are more common: fixed-cost amortization, go-to-market advantages, brand and trust, and market power. Scale effects come with diseconomies of scale too.
What repeatable data-moat playbooks exist?
- Rich Barton (Expedia, Glassdoor, Zillow): make valuable, fragmented, opaque knowledge public, then dominate search and own customer demand.
- Luis von Ahn (Duolingo, reCAPTCHA): a two-sided learning loop in which users get value from doing the labelling.
- Travis May (LiveRamp, Datavant, Shaper Capital): solving data fragmentation industry by industry.
How should founders and investors use this framework?
- Identify which moats you have or can build: control or loops, and which subtype.
- Judge how strong and durable each is, and where it sits in its lifecycle. Moats overlap, many only work above critical mass, and many have upper limits, including replacement by LLMs.
“Not all data moats are created equal”, and knowing the strengths and weaknesses of your own and your rivals’ moats is itself a strategic advantage.
Related
- Full essay: Data and Defensibility, Pivotal, 12 April 2025
- The Economics of Data Businesses (summary)
- How to Price a Data Asset (summary)