Looks, Brains, and Money (summary)
Summary of an essay by Abraham Thomas, published in Pivotal on 11 August 2026. Read the full essay on Pivotal →
Why doesn’t finance use AI more, and what would fix it? Abraham Thomas argues that “the missing ingredient for finance is data quality.” Finance lags software, law, sales and marketing, and even education and healthcare in AI adoption. Many “AI failures” trace back to bad data quality, either in what is sent to the AI or in how its output is evaluated. The essay gives finance professionals concrete practices for three ways LLMs meet data quality: data input to AI, data output from AI, and AI as an evaluator of data.
It is written for non-technical finance professionals on Wall Street and Main Street: CFOs, finance managers, wealth advisors, analysts, underwriters, FP&A, controllers, and family-office allocators. Thomas says the principles generalize to anyone using AI for professional work. It is part two of a series; part one is On Data Quality.
What is “careless in, convincing out”?
“Garbage in, garbage out” still holds, but LLM outputs are “horribly plausible”, so errors are harder to spot. Beyond obvious garbage, LLMs take “lazy or shoddy or noisy inputs” and polish them into confident assertions. Thomas’s name for this: “careless in, convincing out” (CICO). Everything given to a model counts as input and is subject to quality: prompts, files, connectors, apps, history, and skills.
How should finance professionals manage AI inputs?
The core principle: “maintain closer control of fewer but better data inputs.”
| Practice | What it means | Example from the essay |
|---|---|---|
| Curate ruthlessly | Don’t put everything into context; more is almost always worse | — |
| Provide constant grounding | Use model-independent sources of truth throughout; never let the LLM write to them | — |
| Watch for hidden handoffs | Quality degrades where data passes between AI tools unseen | A wealth advisor’s notes pass through microphone → speech-to-text → diarization → compaction → prompt → model |
| Stay current | Stale files and version conflicts confuse models | 2026-budget-v17-final-James-v2-FINAL.xlsx |
| Inspect the raw material | Look at the rawest version of transcripts, files and tables | An AI-generated invite to a call with “Yumi”, who turned out to be “you, me, and [someone]” |
| Beware sycophancy | Models can lead users into low-quality prompts | — |
| Avoid context rot | Long sessions accumulate confusion and persistent errors | A temporary “8% vacancy” assumption that becomes the model’s actual rate by message 30 |
How should finance professionals check AI outputs?
- Ask for grounding (facts, references, citations, logic) at every stage, not just at the end.
- Make it show its work, because models skip steps, over-extrapolate, and invent plausible data.
- Promote dissent: ask for steelman and devil’s-advocate cases. Questions like “Are you sure? Where did you see that?” work well; treat the model like a junior analyst getting a late-night “please fix”.
- Learn the AI tells in financial work: spurious precision, confusing inputs with assumptions and outputs, internal inconsistency and calculation errors, suspiciously clean results, scale and unit errors, and no sense of materiality.
At the higher rungs of the quality ladder (fitness for purpose and business outcome):
- Beware fluency. LLMs “cosplay competence”. Separate analysis from presentation: get plain facts first, polish later.
- Watch for reward-hacking. A persistent chat learns what you like to hear, and board updates drift into “panglossian treacle”.
- Know what you don’t know. Demand ranges and confidence, not point estimates such as a precise runway date.
- Beware flattening. Push past generic, could-be-any-company output (for example “the macro could worsen”) toward risks specific to you.
- Don’t anchor. Output that only repeats what you know is worthless.
- Loop in humans. Their errors are imperfectly correlated with the AI’s, which diversifies risk.
What can go wrong in a real finance task?
Thomas’s case study is a CFO preparing a quarterly board pack with AI help:
| What happens | Failure type |
|---|---|
| The PDF parser skips a line | Hidden handoffs |
| The revenue sheet uses an outdated contract | Version conflict |
| A missing FX rate is invented by the LLM | Grounding |
| “Great! I have everything I need. Would you like me to write the commentary?” | Prompt leading |
| Results get worse with each iteration | Context rot |
| The wrong transcript is used because it was mislabelled | Inspection gap |
| Some calculations are simply wrong | Showing the work |
| Data carries over from the previous board pack | Anchoring |
| The commentary is dense and convincing | Fluency |
| On a closer read, the conclusions are obvious | No dissent |
| Little is specific to the firm | Flattening |
“All of these are failures of the LLM. But more fundamentally, these are failures of data quality. And they can be fixed if you fix data quality.”
Is Thomas pessimistic about AI?
No. He writes that for tasks LLMs are tuned for, such as software engineering, using them “feels like magic”, “the real thing”, and that “LLMs are incredible.” They are tools that can be used badly. His slogan: “Steer, don’t fear!”
Can AI evaluate its own output?
AI evaluation brings scale, speed, cost, and consistency, but it has two failure modes. Self-certification: models are poor at auditing themselves and tend to double down or oscillate when challenged. Untethering: closed input → output → evaluation loops drift away from reality. AI judges also have their own biases, which can be gamed.
What is contamination, and how do you prevent it?
When AI outputs are written back into an organization’s sources of truth (general ledger, CRM, customer metrics, operating handbook), errors “persist and propagate”: later AI tools read bad data, treat it as correct, and write more. The organization becomes “subtly detached from reality.” The remedy is informational discipline: isolate primary sources, quarantine genuinely proprietary data, flag and separate LLM outputs, supervise, and “chain rarely.” Thomas notes the web has already suffered this through Gresham’s law, bad content driving out good, and calls data quality “a ratchet”: the degradation is often unrecoverable.
Related
- Full essay: Looks, Brains, and Money, Pivotal, 11 August 2026
- Part one: On Data Quality (summary)
- Data in the Age of AI (summary)