Skip to content
noteData Before Models

Context failures, not model failures: reading the agent post-mortems

Production studies show agent failures cluster in scaffolding: 17 percent are step repetitions and 14 percent are reasoning-action mismatches, modes invisible to output-only checks. The fix is trajectory evaluation and better context assembly, not a bigger model.

Mezza AI Research

When an AI agent fails in production, the instinct is to blame the model. The failure data says the instinct is usually wrong.

A December 2025 study of agents in production (arXiv, “Measuring Agents in Production”) cataloged how deployed agents actually break. The two largest failure classes: 17.14 percent of failures are step repetitions, where the agent loops on work it already did, and 13.98 percent are reasoning-action mismatches, where the agent’s stated reasoning and its actual tool call disagree. Neither is a wrong answer in the usual sense. Both are failures of the machinery around the model: state, memory, tool wiring, context assembly. Anthropic’s engineering team had already named the pattern in September 2025: most agent failures are context failures, not model failures.

The practical consequence is uncomfortable for anyone whose quality process only checks final outputs. A step-repeating agent can still emit a correct answer, late and at triple the cost. An agent whose reasoning and actions diverge can land on the right number for the wrong reason, which is the kind of correctness that fails the next time conditions shift. Output checks see none of this. The study’s authors make the point directly: these failure modes are invisible to final-output evaluation.

Evaluate the trajectory, not the answer

The discipline that emerged through 2025 and 2026 evaluates the full trajectory: was each tool choice correct, were the arguments valid, how many steps did the task take, what did it cost, did every step comply with policy. Teams run deterministic mocks in continuous integration and score live trajectories with LLM judges. A tooling category formed around this need (Galileo, Arize, Langfuse, LangSmith, Maxim), which is itself evidence the problem is real and general.

For a financial institution the trajectory is more than a debugging aid. It is the audit record. A regulator asking why the system flagged one counterparty and not another is asking a trajectory question. Firms that log and evaluate trajectories from day one get compliance evidence as a byproduct of their quality process. Firms that only check outputs will rebuild their evaluation stack the first time an examiner asks how a number was produced.

What buyers should take from this

Three questions sort vendors quickly.

First: what fraction of your failures are context failures, and how do you know? A vendor who cannot answer has never measured. A vendor who answers with output-accuracy numbers has measured the wrong thing.

Second: show me a trajectory. Not a demo answer, the full record of one production-grade task: the plan, the tool calls, the retrievals, the dead ends, the stop condition. The record either exists in inspectable form or it does not.

Third: where do humans sit in the loop? The settled best practice is that agents recommend and humans approve high-impact actions, with the approval itself logged (Elementum AI, 2025). An approval gate without a trajectory behind it asks the human to approve blind, which protects no one.

The encouraging reading of the failure data is that the hard problems are engineering problems. Step repetition has known fixes in memory and state design. Reasoning-action mismatch yields to better tool definitions and trajectory tests. Context failures yield to context engineering. None of this waits on a research breakthrough. It waits on someone doing the unglamorous layer work well, which is precisely the work most pilots skip.

The model will keep getting credit for the successes and blame for the failures. The record says both belong mostly to the layer around it.

Mezza AI Research.

Works cited

  1. arXiv, Measuring Agents in Production, 2025-12
  2. Anthropic, Effective Context Engineering for AI Agents, 2025-09
  3. Andrii Furmanets, AI Agents in 2026: Tools, Memory, Evals, and Guardrails, 2026
  4. Elementum AI, Human-in-the-Loop Agentic AI, 2025

Topics: evaluation, agent failures, trajectories, observability