Data before models: what actually decides whether agentic AI scales
Most financial firms are piloting AI agents and few have them in production. The blocker is not model capability. It is fragmented data, missing semantics, and absent lineage. The firms that scale agentic AI are the ones that build governed data products first and treat the model as the replaceable part.
The model is not the bottleneck. Across the analyst record of 2025 and 2026, one finding repeats: whether agentic AI reaches production is decided in the data layer. McKinsey’s work on scaling agentic AI states it directly: data infrastructure, not the model, determines whether agentic systems scale, with roughly 80 percent of companies citing data limitations as the roadblock. Fortune ran the same conclusion as a headline in April 2026. The deployment gap is stark in finance: at the end of 2025, 68 percent of asset management and private equity firms were piloting AI agents while 24 percent were deploying them (KPMG AI Quarterly Pulse, Q4 2025). The pilots are not waiting on a better model.
Update, September 2026: the Q2 2026 edition of the same survey reports 52 percent of asset management and private equity firms piloting AI agents and 19 percent deploying them (KPMG AI Quarterly Pulse Survey, Q2 2026).
This essay lays out the data-before-models thesis for financial institutions: why the gap exists, what the data layer must contain before agents can be trusted on it, and why the build order (data products first, agents second) is the only sequence that survives contact with production. The thesis is not contrarian. Clearwater Analytics compressed it into a slogan: AI is only as powerful as the data underneath it. We agree with the slogan. The argument deserves more than one line.
Why do agent pilots stall?
Because pilots run on curated data and production runs on the real thing. A demonstration agent answers questions over a clean folder of documents someone prepared for it. A production agent must operate over the actual estate: positions in one system, contracts in another, reference data in a third, definitions that disagree, entities that appear under four names. The 2026 vendor consensus in finance is blunt about this. Agentic deployments fail not because models are immature but because financial data is fragmented, with no agent-ready intelligence tier above it (Techment, 2026).
There is a second, quieter reason. Anthropic’s engineering team observed in September 2025 that most agent failures are context failures, not model failures. The agent did not reason badly. It reasoned correctly over the wrong inputs: a stale price, the unsigned draft instead of the executed contract, revenue under last year’s definition. Pilots hide this failure mode because someone hand-fed the context. Production exposes it, because in production the data layer is the feeding mechanism. If that layer cannot say which document is authoritative or which definition is current, no model can compensate downstream.
Throwing a bigger context window at the problem fails for a measured reason. Anthropic describes context rot: recall degrades as the window grows, because attention is a finite budget spread across everything included. More raw data in the prompt makes retrieval worse, not better. The answer is curation upstream, which is data-layer work.
What does an agent-ready data layer contain?
The clearest taxonomy in the current literature separates three artifacts that get conflated under “semantics” (Context and Chaos, January 2026; Atlan, 2026). Each answers a different question, and agents need all three.
A semantic layer standardizes metrics: define revenue, exposure, and NAV once, so every consumer computes the same number. Gartner elevated the semantic layer to essential infrastructure in its 2025 Hype Cycle for analytics, with analysts now describing it as a control plane for enterprise AI. Without it, an agent cannot calculate.
An ontology encodes domain knowledge: the concepts, relationships, and constraints that give financial data meaning. That a bond has an issuer, that an issuer has a parent, that a guarantee transfers exposure up the chain. Without it, an agent cannot reason.
Lineage and context graphs record where data came from, how it was transformed, and why past decisions were allowed. Decube’s 2026 framing calls lineage essential for deploying agentic AI with confidence and auditability: it shows exactly which data contributed to a decision and how. Without it, an agent cannot be trusted, in the literal sense that no auditor can verify its output.
A fourth element completes the finance version of the list: policy encoded as deterministic rules. The 2026 vendor framing in financial services is explicit that an agent-ready layer includes the firm’s actual financial policies written as rules a machine enforces, not as guidance a model is prompted to follow (Techment, 2026). Concentration limits, restricted lists, approval thresholds, and disclosure duties are not suggestions an agent should weigh. They are constraints the layer must enforce before any output leaves the system. Probabilistic reasoning on top, deterministic boundaries underneath. A firm that has never written its policies down in machine-checkable form discovers, usually mid-project, that this codification is the longest task on the plan.
The composition matters more than any single piece. An ontology without metric governance leaves agents unable to calculate. A semantic layer without domain knowledge leaves them unable to reason. A knowledge graph without governance leaves them unable to enforce trust. The Context and Chaos analysis calls the whole thing a knowledge architecture problem, needing librarian and taxonomist skills alongside data engineering. That is an unusual hiring profile for a fund. It is a normal one for a finance-native partner.
What does the build path look like?
MIT Technology Review’s April 2026 account of firms rebuilding their data stacks for AI describes a consistent sequence. Consolidate data into open formats, governed with precision. Establish what it calls a critical intermediary between raw data and operational agents: a catalog providing discovery, access control, and business semantics. Codify organizational metrics and context so the AI understands business constraints. Then, and only then, point agents at it. The firms that succeed treat this as a repeatable data-product methodology rather than a string of isolated pilots.
One line from that reporting deserves a financial-services underline: agents need the same governance rigor applied to human employees. Entitlements, action constraints, monitoring, and accountability. A new analyst does not get the whole data estate on day one. Neither should an agent. This is only enforceable if the data layer knows who is allowed to see what, which is, again, infrastructure that exists before any model arrives.
The build order has a strategic consequence worth stating plainly:
The model is the replaceable part. The governed data layer is the asset.
Models improve quarterly and get swapped. The semantic definitions, ontology, lineage, and encoded policies are specific to the firm, compound in value, and move with you across model generations. Investment in the data layer survives every model migration. Investment in prompt-level workarounds for bad data survives none of them.
Does the data-first sequence actually pay?
The efficiency evidence says the prize for getting this right is large. McKinsey estimates asset managers could capture 25 to 40 percent of their total cost base through AI-enabled transformation. Adoption pressure is real: by late 2025 most hedge funds reported some form of generative AI use. The constraint on converting usage into production value is the one this essay describes. KPMG estimated 50 billion dollars of global agentic-AI spend in 2025. A large share of that spend will stall at the pilot stage for data reasons the buyers could have diagnosed in advance.
The diagnosis is straightforward. Before any agent project, ask four questions of the target workflow. Is there one governed definition of every metric the agent will compute? Can the firm say, for any record, where it came from and what touched it? Are the entities resolved, with the same counterparty one thing everywhere? Are the policies the agent must obey written down as rules a machine can enforce, rather than as institutional folklore? Four yes answers mean the workflow is agent-ready. Any no answer is the project plan.
What this means for buyers
Three practical conclusions follow from the evidence.
First, sequence your spend. Money put into semantics, lineage, and entity resolution is not a delay to the AI program. It is the AI program. The agent on top is the cheap part, increasingly so as models commoditize.
Second, interrogate vendors about the layer, not the demo. Any platform can answer questions over clean documents. Ask how it handles disagreeing definitions, how it records lineage, and what happens when the same deal appears in four systems. The answers separate intelligence layers from chat interfaces.
Third, treat data work as the durable differentiator. Your competitors can buy the same models you can. They cannot buy your governed data layer, because it has to be built on your estate, your definitions, and your constraints. In a market where model capability is evenly distributed, the data layer is where advantage accumulates.
The slogan version fits in one line: data before models, because the models are rented and the data layer is owned.
Mezza AI Research. All figures cited are third-party research findings, attributed inline. None are Mezza claims.
Works cited
- McKinsey, Building the Foundations for Agentic AI at Scale, 2025
- McKinsey, How AI Could Reshape the Economics of the Asset Management Industry, 2025
- KPMG, AI Quarterly Pulse Survey: Asset Management and Private Equity, Q4 2025
- KPMG, AI Quarterly Pulse Survey: Asset Management and Private Equity, Q2 2026
- Fortune, Why Your Data Infrastructure Will Determine Whether Agentic AI Scales, 2026-04
- MIT Technology Review, Rebuilding the Data Stack for AI, 2026-04
- Gartner (via Ontoforce), Semantic Technologies Take Center Stage, 2025-04
- Atlan, Semantic Layers, Ontologies, and Context Graphs for AI, 2026
- Clearwater Analytics, The Investment Platform Built for AI (corporate positioning), 2025
- Context and Chaos, Ontologies, Context Graphs, and Semantic Layers, 2026-01
- Anthropic, Effective Context Engineering for AI Agents, 2025-09
- Techment, AI in Financial Workflows: Reshaping Finance in 2026, 2026
- Decube, Data Lineage: The Foundation of AI-Ready Data, 2026
