A loop, not a pipeline: the one diagram that explains agentic AI
Classic RAG retrieves once and answers. Agentic RAG runs a control loop that retrieves, reasons, decides, and acts until a stop condition is met. The loop is what lets a system notice its own evidence gaps, which is the property financial work actually requires.
If you want one diagram that separates the current generation of AI systems from the last, draw a loop next to a pipe.
The pipe is classic retrieval-augmented generation. A query comes in. The system retrieves the top matching passages, stuffs them into the context window, and generates an answer. One pass, one direction, done. It works when the question is simple and the first retrieval happens to fetch the right evidence. It fails silently when the first retrieval misses, because nothing in a pipe can notice a gap.
The loop is agentic RAG. Towards Data Science put the cycle plainly in March 2026: retrieve, reason, decide, repeat until a stop condition is met. The agent reads what came back, judges whether the evidence supports an answer, and acts on that judgment. It can re-query with a sharper question. It can switch tools, moving from document search to a database query to a verification step. It can decompose one hard question into several answerable ones. Graph-based variants extend the loop to multi-hop reasoning across connected entities (arXiv, June 2025).
Why the difference matters in finance
Financial questions are almost never one-retrieval questions. “How exposed are we to this counterparty” touches positions, legal entities, collateral agreements, and yesterday’s news. A pipe answers with whatever the first search returned. A loop notices that the entity appears under three names, resolves them, pulls the collateral terms, and only then answers. The loop behaves the way an analyst behaves. That is not a metaphor. It is the same control structure: gather, assess sufficiency, gather again, conclude.
The loop also produces something a pipe cannot: an inspectable trajectory. Each iteration is a recorded decision about what to fetch and why. For a compliance team, that record is the difference between an answer and an auditable answer.
At production scale the loop becomes an orchestra. Anthropic’s multi-agent research system (June 2025) showed the pattern: a lead agent decomposes the question and spawns parallel specialists, each running its own loop, with a dedicated citation agent attributing sources at the end. The multi-agent version outperformed a strong single agent by 90.2 percent on Anthropic’s internal evaluation, at roughly 15 times the token cost. Both numbers matter. The architecture wins on quality. It also makes cost a managed variable rather than a constant.
The honest tradeoff
Loops are less predictable than pipes. A pipe has fixed latency and fixed cost. A loop’s latency and cost are distributions with tails, because the agent decides at runtime how many iterations the question deserves. The engineering answer is budgets and stop conditions, set deliberately per workflow. A tearsheet can afford a long loop. A pre-trade check cannot. Deciding where each workflow sits on that spectrum is design work, not a default.
This is why we keep insisting the intelligence layer is the asset. The loop needs governed data to retrieve from, memory to avoid re-fetching what it already knows, and provenance on every iteration. Strip those away and the loop is just an expensive way to be wrong several times before answering.
One drawing, then. Raw data at the bottom. Decisions at the top. In between, not an arrow but a circle. If a vendor’s architecture diagram shows a straight line from your documents to an answer, you are looking at a pipe with better marketing.
Mezza AI Research.
Works cited
- Towards Data Science, Agentic RAG vs Classic RAG: From a Pipeline to a Control Loop, 2026-03
- arXiv, Reasoning RAG via System 1 or System 2, 2025-06
- Anthropic, How We Built Our Multi-Agent Research System, 2025-06
- Mem0, Agentic RAG vs Traditional RAG: Complete Guide, 2025
