The Bounded Autonomy Ladder: how much should a financial AI agent be allowed to do?
The Bounded Autonomy Ladder is a five-level maturity model for agentic AI in regulated finance: Retrieve, Draft, Propose, Approve-to-Act, and Mandated Autonomy. Each level must be earned by the controls beneath it. The ladder converts the vague question of whether to trust AI into the answerable question of which control is missing for the next level.
How much should a financial AI agent be allowed to do? The wrong answer is a number. The right answer is a ladder. Autonomy in regulated finance is not a setting you choose. It is a position you earn, one control at a time. This essay defines the Bounded Autonomy Ladder, a five-level maturity model for agentic AI in financial institutions, with the gating control named at every rung.
The Bounded Autonomy Ladder is a five-level maturity model for agentic AI in regulated finance, in which each level of autonomy (Retrieve, Draft, Propose, Approve-to-Act, Mandated Autonomy) must be earned by the controls beneath it before a firm climbs to the next.
The ladder exists because the industry’s current question is unanswerable. “Do we trust AI?” admits no useful reply. “What level are we at, and which control is missing for the next one?” admits a precise reply, a budget, and next steps. Compliance teams can work with the second question. The ladder is how you ask it.
Why does autonomy need a maturity model?
Because trust, not capability, is what gates deployment. The CFA Institute found lack of explainability the second most cited barrier to AI adoption among investment professionals (via Planful, 2025). PwC’s 2025 Global Investor Survey of 1,074 professionals found investors expect firms to disclose AI governance, model validation, and security controls, not just outcomes. A Billtrust study of 500 finance professionals (November 2025) found 82 percent concerned about AI misuse. Meanwhile the regulatory posture hardened: the SEC and OCC are examining AI governance directly, MAS proposed formal AI Risk Management Guidelines in November 2025 with a 12-month transition, and examiners treat missing decision traces as a books-and-records problem (Galileo, 2025; Banking Exchange, 2025).
Against that backdrop, the industry settled on a phrase: bounded autonomy. Agents assist, propose, and prepare; humans approve high-impact actions; everything is logged (Neurons Lab, 2026; Elementum AI, 2025). The phrase is right. What it lacks is resolution. Bounded how, at what level, with which controls? The ladder supplies the resolution.
The five levels
L0: Retrieve
The agent answers questions over governed data and cites every source. It takes no action and produces no artifact beyond the cited answer. Humans do all of the work; the agent compresses the search.
Gating control: grounded citation. Every claim traces to a source a human can open. Retrieval logs record what was fetched, from where, with what transformation. This is not decoration. Agentic retrieval with enforced citations is the mechanism that narrows the governance gap at the base of the ladder. Hebbia’s sentence-level citations and Kensho’s grounding work show the market converging on this floor. A firm that cannot ground answers in sources has no business on any higher rung.
L1: Draft
The agent produces work artifacts for human revision: tearsheets, earnings previews, reconciliation summaries, model documentation, first-draft memos. The human owns every artifact that leaves the desk. This is where most credible production deployments in finance sit today. S&P’s Kensho plugin builds tearsheets and earnings previews from Capital IQ data (February 2026). The broad pattern across institutions is agents working research backlogs, reconciliations, and reporting drafts (Neurons Lab, 2026).
Gating control: provenance on the artifact. A draft is only safe to edit if the reviewer can see which inputs produced which passages. Lineage from source to sentence turns review from re-research into verification. Without it, checking the draft costs as much as writing it, which deletes the economics of the rung.
L2: Propose
The agent recommends a specific action and presents its evidence: rebalance candidates, a flagged counterparty, an exception worth escalating. A human decides. The finance literature’s synthesizing orchestrator, weighing specialist signals and reconciling conflicts, lives at this level. So does the discipline of confidence signaling, where research on production agent interfaces found simple high/low indicators outperform numeric scores because reviewers decide faster (Fuselab Creative, 2025).
Gating control: the decision record. Every proposal is logged with its evidence, its alternatives, and how the human disposed of it. Accepted or rejected, the record persists. This rung is where the audit trail becomes a decision trail, the thing SR 11-7-style model risk review actually wants to read: not just what the system said, but what was done about it and on what basis.
L3: Approve-to-Act
The agent executes a multi-step workflow after explicit human approval: refresh the model, rerun the reconciliation, prepare and stage the order package. Approval covers the plan; the agent carries out the steps. Rogo’s Felix, executing multi-step financial processes across origination and portfolio work, and BlackRock’s supervised Aladdin Copilot architecture are the market’s current expressions of this rung.
Gating control: trajectory evaluation. Before a firm lets an agent execute, it must be able to evaluate full trajectories, not final outputs. The production failure data explains why: a December 2025 study found 17.14 percent of agent failures are step repetitions and 13.98 percent are reasoning-action mismatches, modes invisible to output-only checks (arXiv, “Measuring Agents in Production”). At L3 a bad trajectory is no longer a bad draft. It is actions taken in the world. The firm needs trajectory tests in CI, runtime monitoring, and kill-path controls before the first approval is granted.
L4: Mandated Autonomy
The agent acts without per-action approval, inside deterministic, pre-encoded policy limits: thresholds, entitlements, instrument whitelists, exposure caps, escalation triggers. Humans set the mandate, monitor aggregate behavior, and review on a cycle. This is the rung agentic payment rails are built for. Google’s AP2 protocol, with its signed Intent, Cart, and Payment mandates carried as verifiable credentials (September 2025), is an early formalization of machine-checkable mandate boundaries.
Gating control: encoded policy plus runtime enforcement plus agent identity. The recurring four-part framework for compliant agent deployment applies in full here: agent identity, runtime enforcement, comprehensive auditing, and lineage (Promethium, 2026). The policies must be rules a machine enforces at runtime, not guidance a model is prompted to respect. The agent must have an identity and entitlements governed like an employee’s. The trade-off is honest: L4 narrows what the agent may do in exchange for not asking permission each time. A mandate too wide to enforce deterministically is not a mandate. It is hope.
How should a firm use the ladder?
Locate, then name the gap. Map each AI workflow to its current rung, honestly. Most firms discover they run a portfolio of rungs: L1 in research, L0 in compliance queries, an ambitious pilot reaching for L3 without the trajectory evaluation that gates it. The ladder’s first service is making that mismatch visible.
Climb one rung at a time, per workflow. The ladder is per-workflow, not per-firm. A reconciliation process with clean lineage may earn L3 while research drafting stays at L1. Skipping rungs is how firms end up with what the disclosure record shows in the aggregate: 72 percent of the S&P 500 disclosed a material AI risk in 2025 filings (The Conference Board, 2025). Ambition outran controls.
Buy the controls, not the rung. Vendors sell autonomy. The ladder says autonomy is downstream of controls, so procurement questions should target the gating control directly. Show me the citation mechanism. Show me artifact lineage. Show me a decision record. Show me trajectory evaluation. Show me runtime policy enforcement. A vendor strong on rung language and weak on control evidence is selling a position they have not earned.
Let the ladder set the meeting agenda. The recurring institutional deadlock is a business side that wants agents and a risk side that wants assurances, with no shared vocabulary between them. The ladder gives both sides the same object: this workflow is at L1, the business case wants L3, and here are the two named controls in between. That is a fundable conversation.
What the ladder is not
It is not a capability benchmark. A model can be brilliant and belong at L0 because the data beneath it lacks lineage. It is not a one-way ratchet. Workflows should descend a rung when a control degrades, when data sources change, or when monitoring flags drift. It is not a promise that L4 is the destination for everything. Plenty of financial workflows should live permanently at L2, because the cost of a wrong autonomous action exceeds the cost of a human glance forever. The ladder’s claim is narrower and more useful: wherever a workflow ought to sit, the way up is named controls, in order, with evidence.
The question to retire is whether to trust AI. The question to adopt is one rung more specific.
Mezza AI Research. All figures cited are third-party findings, attributed inline. None are Mezza performance claims. The Bounded Autonomy Ladder is a Mezza AI Research framework; cite it with a link to this page.
Works cited
- Galileo, AI Agent Compliance and Governance, 2025
- Billtrust, AI in Finance Survey of 500 Finance Professionals, 2025-11
- Banking Exchange, Compliance for AI Agents in Financial Services, 2025
- Promethium, AI Agent Data Governance: The Enterprise Playbook, 2026
- Elementum AI, Human-in-the-Loop Agentic AI, 2025
- arXiv, Measuring Agents in Production, 2025-12
- PwC, Global Investor Survey 2025, 2025
- Planful, Why You Need Explainable AI in Finance, 2025
- Neurons Lab, Agentic AI in Financial Services: A Research Roundup, 2026
- Kensho, S&P Global Brings Financial Skills to AI Agents, 2026-02
- lowtouch.ai, How BlackRock's Aladdin Copilot Uses Agentic Architecture, 2025
- Google Cloud Blog, Announcing Agent Payments Protocol (AP2), 2025-09
- The Conference Board, AI Risk Disclosures in the S&P 500, 2025-10
- Fuselab Creative, AI Agent UX Design Patterns, 2025-08
