On Data Quality (summary)
Summary of an essay by Abraham Thomas, published in Pivotal on 27 June 2026. Read the full essay on Pivotal →
What is data quality? Abraham Thomas’s answer: “data has no innate quality”; quality exists only relative to what the data is used for. “Data quality is that which increases data value.” He organizes it as a ladder of four ordered, dependent levels: granular (unit) quality, aggregate (corpus) quality, fitness for purpose, and business-outcome quality. “Quality is a ladder. The lower rungs enable the higher ones; the higher rungs justify the lower ones.”
This is part one of a two-part series. Part two, Looks, Brains, and Money, applies the framework to AI in finance.
Why are the standard definitions of data quality not enough?
- ISO 8000 defines quality data as data that meets its stated requirements: “perfectly accurate and completely useless.”
- ISO 25012 defines quality through 15 attributes such as accuracy, completeness and consistency: correct, but incomplete.
Thomas starts instead from his earlier argument in How to Price a Data Asset: data has no intrinsic value; its value is the value of what can be done with it. Data quality is whatever increases that value, so it can only be judged against a use.
What are the four levels of data quality?
| Level | What it covers | Example attributes | Questions it answers |
|---|---|---|---|
| 1. Granular (unit-level) | A single record, sentence, Q&A pair or labelled example, judged on its own | Accuracy, precision, recency, well-formedness, internal consistency, plausibility, provenance, interpretability, confidence | Is it true, usable, current, coherent? |
| 2. Aggregate (corpus-level) | The dataset as a whole | Coverage, deduplication, granularity, representativeness and balance, cross-record and label consistency, distributions, sufficiency, continuity, joinability, drift | Is it all there, clean, representative of the world, stable over time and space? |
| 3. Fitness for purpose | The fit between the data and a specific application | Informational fit (relevance, adequacy, sufficiency, necessity); operational fit (availability, licensing and compliance, interoperability, risk/reward calibration) | Does it answer the questions you have, and can you use it effectively? |
| 4. Business outcome | Whether using the data creates value | Adoption, influence on decisions, change in actions and outcomes, attribution, materiality, ROI, timeliness, durability, risk | Was it used, did it change anything, and was the change worth it? |
The levels exist simultaneously, and much disagreement about data quality comes from people talking about different levels. Thomas’s image is the parable of the blind men and the elephant.
How does one example move up the ladder?
The essay follows a company’s revenue data up all four rungs:
- Granular: booking the right revenue number means reading contract terms, renewals and discounts correctly, and even then, gross versus net reporting for a marketplace is a judgement call.
- Aggregate: a revenue history can be wrong as a whole if definitions changed midway, entries are missing or double-counted, or totals don’t reconcile. A perfect snapshot of today’s customers may still be unrepresentative for forecasting expansion.
- Fitness for purpose: books closed accurately a few days after month-end are high quality for an auditor but too late for a CEO making decisions mid-month, and the detail that serves finance overwhelms a board.
- Business outcome: perfect revenue data used to redesign sales bonuses can still backfire if the sales team games the new formula. “At the highest level, data quality is not about the data.”
What are the common failure modes?
- Failure to launch: perfecting the lower rungs (granular, aggregate, purpose) while delivering no business value. It is “astonishingly common”, because the lower rungs are tangible, measurable and legible.
- Failure to ground: skipping the lower rungs to chase business results. It can work briefly with a clear target and fast feedback, but is rarely sustainable; “foundations matter.”
- Provenance as proof: the “maybe-okay” shortcut of borrowing quality from a trusted source, which lets you spend less on checking lower-level quality. Trust in a source is earned over time and lasts only as long as the data keeps working.
The practical test: on the lower rungs, ask whether you are neglecting the business use case; on the higher rungs, ask whether you are neglecting foundational hygiene.
Related
- Full essay: On Data Quality, Pivotal, 27 June 2026
- Part two: Looks, Brains, and Money, on data quality and AI in finance
- How to Price a Data Asset (summary)