AI Agents Executing Real Bank Transactions vs. Surfacing Insights
Banks must rethink governance as AI agents execute transactions instead of just recommending them.

Agentic AI in banks doesn't just flag a suspicious transaction or suggest a next step. It moves the money, opens the case, settles the trade. That shift, from telling a human what to do to doing it directly, is forcing banks to rethink what governance even means once there's no one left in the loop to say yes.
For most of the last twenty years, the AI sitting inside a bank's systems has been an advisor, not an actor. A model flags a transaction as likely fraud. A credit score gets calculated. A dashboard surfaces a recommended action. In every case, a human stands between the machine's output and any real-world consequence: someone reads the flag, checks the score, decides whether to act. Agentic AI removes that step. The system takes an objective, works out a plan, and carries it out across whatever systems it needs to touch, often with little sign-off along the way. Creatio, the CRM and process-automation vendor, draws the line cleanly: traditional AI automates tasks or hands over insight, while agentic AI runs banking processes start to finish. The IMF put a sharper point on it in an April 2026 note (Davidovic and Tourpe, IMF Notes 2026/004): in payments specifically, agentic AI risks turning the basic unit of a transaction from a human-initiated instruction into an agent-mediated decision.
That's a difference beyond a faster chatbot. It's a different relationship between person and machine, one where the agent acts and the human's job becomes governing the acting, not doing it.
What agentic AI is already doing inside banks, and where execution ends and insight begins
Most of what banks call "AI" right now still lives on the insight side of that line. Worth being honest about the gap between "using agentic AI" and actually letting it run.
JPMorgan's COiN system reviews commercial loan agreements, work that used to eat more than 360,000 lawyer-hours a year. It reads documents and flags terms. It doesn't execute anything. Bank of America's Erica handled around 2 million customer interactions a day as of 2024 and crossed 3 billion lifetime interactions by August 2025, walking people through balances and spending patterns. Erica recommends. It doesn't move money on its own. Axis Bank's voice assistant handles more than 100,000 voice requests daily, helping customers understand loan options or installment schedules. Same pattern: surfacing, not acting.
Then there's a middle band, where AI starts touching workflow even though a human still pulls the final trigger. Scotiabank's in-house tool, AIDox, reads incoming client emails in commercial banking, works out what's being asked, routes it to the right team, and opens a case in Scotiabank's own system. That's a step past pure insight: the AI decides where things go, not just what they mean. A human still executes the transaction. Finzly went further in October 2025 with its "Agentic Galaxy," a set of AI agents built into Finzly's existing tech stack for banks and credit unions, covering payments, foreign exchange, and virtual accounts. Finzly calls this execution infrastructure, not a dashboard, and the framing matters.
Full execution is where the actual category shift happens. A global standard-setting body cited in the same IMF note found that generative AI systems can hold precautionary liquidity buffers, prioritize urgent payments, and weigh liquidity cost against settlement delay, without being specially trained to do any of it. The IMF describes agentic systems reallocating funds on their own, based on preset parameters and live market conditions. Cross-border payment orchestration is the clearest case: one agent handling initiation, routing across correspondent banks, compliance checks, settlement monitoring, and exception handling, start to finish, with no human touching a single step.
So why does this spectrum matter to anyone trying to size up the risk? Because most of what's deployed today still sits comfortably on the insight end. A 2025 MIT Technology Review survey of 250 banking executives found 70% already using agentic AI in some form, through live deployments or pilots. Only 16% had anything actually live. That gap, between banks talking about agentic AI and banks letting it execute, is where the industry actually sits right now, and it's a wide one.
Why the execution layer creates governance problems that the insight layer never had
When AI only surfaces something, a human decision and a human action still stand between the machine and any consequence. The error surface stays narrow. Someone can catch a bad recommendation before it does damage, and even if they don't, the damage is usually reversible: a wrong transfer gets clawed back, a bad tip gets ignored next time.
Execution collapses that gap. The agent is the decision and the action, happening in the same instant. A misconfiguration, a bug, or an adversarial prompt can produce a financial outcome before anyone notices it, let alone approves it.
The IMF's April 2026 note names a core tension: AI agent behavior is adaptive, while payment infrastructure demands determinism. Settlement is supposed to be final. Agents, by design, are supposed to adapt. Put those two properties in the same pipe, and once a transaction reaches legal finality, undoing it may not be possible at all. That's categorically different from a misrouted email or a recommendation a banker chose to ignore. The core risk the IMF names is simple: letting adaptive systems make irreversible payments without checks built in first.
There's a second problem, and it's systemic rather than individual. What happens when multiple banks run agents trained on similar models, reading similar market signals? The IMF flags the possibility that agents could act in correlated ways, amplifying stress across payments infrastructure in ways the system was never designed to absorb. That's a new category of risk, tied specifically to systems that act rather than advise.
And here's the part worth sitting with: even the safeguard meant to fix all this carries its own cost. Analysis cited in the IMF's note points out that requiring human approval for every payment can itself introduce delay, and that delay can undermine the very controls the approval step was supposed to protect. The fix for one harm quietly creates another. Insight AI never faced this trade-off, since a human was always going to review the output anyway. With execution AI, what matters is whether the model's advice was sound. It's whether the agent's own action was sound, and whether anyone can say who's accountable when it wasn't, or whether it can be undone at all.
The architecture regulators and standard-setters are converging on for agents that execute
A few institutions have already sketched what responsible execution architecture looks like, and the convergence in their answers is hard to miss.
The IMF's own framework, from the same April 2026 note, splits the problem into three layers. Layer one is intent formation and orchestration: where the agent reasons, plans, and adapts, and where probabilistic behavior actually belongs. Layer two is authorization and control: rule-based, checking the agent's proposed action against fixed mandates before money moves. Layer three is settlement: deterministic, final, legally binding, with no room for probabilistic behavior at all. The logic is simple once it's laid out: push the adaptive reasoning as far upstream as possible, and keep everything that requires legal certainty completely deterministic.
The Financial Stability Board's consultation report, published June 10, 2026, lays out twelve practices for responsible AI adoption. It accepts something practical: watching every individual agent decision in real time stops being realistic once agents operate at scale. Instead, for agents executing transactions involving customer funds, the FSB recommends human approval or dual authorization above a set value threshold, comprehensive audit trails covering agent activity. It also recommends something that sounds almost circular: using AI to monitor other AI, a category the FSB now formally recognizes.
IOSCO's Supervisory Toolkit for AI Use in Capital Markets, finalized May 25, 2026, approaches the same idea from a different angle: some member authorities are experimenting with "AI as a judge," systems built specifically to assess and oversee other AI systems. Gartner projects these "guardian agents" could capture 10% to 15% of the entire agentic AI market by 2030. Coinbase's agent wallet, launched February 2026, is a working example already shipped, with programmable guardrails built directly into the wallet itself.
One more thing worth flagging. SR 26-2, the interagency model risk guidance issued in April 2026 to replace the older SR 11-7, explicitly excludes generative and agentic AI from its scope, calling both "novel and rapidly evolving." Translation: the rulebook banks have leaned on for model governance doesn't cover this yet. Which means the design choices individual institutions make right now matter more than usual. Regulators haven't finished writing the rules those choices will eventually be judged against.
What banks that are deploying execution agents are actually doing to stay in control
The pattern among early movers is a staged rollout, one that starts with the workflows where a mistake is visible and cheap to fix.
Exception handling, invoice matching, collections support, dispute triage: these are agentic AI's proving grounds right now. The agent works on the workflow around the money, not the money itself. CFO sentiment tracked in May 2025 shows how fast this shifted: 85% of CFOs surveyed said they had no plans to deploy agentic AI at all. By July 2025, follow-up data showed it already moving into early test runs at a small but growing number of enterprises, under tight guardrails. Two months, a near-total reversal in stated intent. That's the pace this space moves at.
Scotiabank's AIDox shows the staging logic in practice. The agent decides where an email-based request should go and opens the case. A human team still executes the transaction. The control layer deliberately keeps "the agent can act on the workflow" separate from "the agent can move funds," and that separation is the whole point.
A handful of mechanisms keep showing up across these deployments:
- Mandate-based authorization, where agents work within explicitly bounded instructions rather than open-ended goals
- Transaction thresholds that trigger human or dual sign-off above a set value, consistent with the oversight approach regulators have begun articulating
- Kill switches, built to interrupt an agent before an action settles, since reversing something after settlement often isn't an option
- Full audit trails covering not just what the agent did, but the reasoning and routing steps that led there
- Restricted system access, so an agent is only credentialed for the specific systems and transaction types inside its defined scope
The payoff for building all this out is real. PwC analysis found agents cutting cycle times by up to 80% in purchase order processing and matching, while actually improving audit trails rather than weakening them. McKinsey found agentic AI cutting manual workload in banking operations by 30% to 50%. Source-of-wealth verification, a compliance-heavy process that used to take 10 days, has been brought down to about an hour.
The lesson underneath all of it is simple, even if it's easy to skip: deploying an agent and granting it full execution authority are two separate decisions. Banks that treat them as one are skipping the control layer the IMF and FSB both treat as mandatory, not optional. The distinction carries real weight. It's the whole ballgame.
What deploying execution agents on existing banking rails (rather than replacing infrastructure) changes about the governance calculus
Easy detail to miss: most of the execution deployments described above don't run on new infrastructure at all. AIDox plugs into Scotiabank's existing case-management system. Finzly's Agentic Galaxy sits inside Finzly's existing tech stack rather than replacing it.
Why does that matter so much? Because the authorization and settlement layers, the parts of the IMF's three-layer model carrying the highest stakes, already exist inside these institutions, complete with their existing controls, logging, and compliance wiring. The agent slots into layer one, intent and orchestration, without ever touching layers two and three. Nothing about settlement changes. Nothing about how a transaction gets authorized changes. Only the reasoning that decides what to propose changes.
Compare that to the alternative: building entirely new execution rails purpose-built for agents. That path means re-proving every compliance and audit property the old rails already had, from scratch, for a system regulators haven't seen before. It also opens a version of the correlated-behavior risk mentioned earlier, but at the infrastructure level. If several institutions migrate onto similar agent-native rails around the same time, any shared flaw or blind spot gets replicated across the industry at once.
The interface layer tells the same story from a different angle. When a customer talks to Axis Bank's Aha! or asks Erica a question, the agent is reaching into existing core banking systems through an API. It isn't replacing those systems. Aha! handling over 100,000 voice requests a day works precisely because the interface changed while the deterministic rails underneath, and their audit trails, stayed exactly where they were.
For anyone deciding whether to greenlight an execution agent, here's the question worth asking first: can what the agent did be audited after the fact? That question gets a lot easier to answer when an agent's actions already flow through systems built to produce regulator-grade audit trails, instead of through something new and unproven.
The accountability standard banks should hold execution AI to, and why it is achievable now
The shift from insight to execution is real and already underway, with 70% of banking leaders in deployments or pilots according to the MIT Technology Review's 2025 survey. That pace is reason to keep moving. It's also reason to be exact about what "in control" actually means, rather than treating the phrase as a given.
Four properties, drawn from the IMF, FSB, and IOSCO frameworks above, separate an execution agent worth trusting from one that isn't. Architectural separation comes first: adaptive reasoning upstream, deterministic authorization and settlement downstream, exactly as the IMF's three-layer model lays out. Second is a bounded mandate, meaning the agent's execution authority is scoped and credentialed, never open-ended. Third is a full audit trail, one that captures not just outcomes but every step of an agent's reasoning, logged in a form that holds up to regulatory review, the FSB's stated bar for any agent touching customer funds. Fourth is tiered human oversight, which isn't a blunt choice between approving every single action and full autonomy, but thresholds, kill switches, and AI systems built to monitor other AI, the point where the FSB and IOSCO both land.
None of this should get sold as a differentiator. In a regulated industry, an execution agent needs to demonstrate these properties, including baseline certifications like SOC 2, just to count as a serious option to begin with. It's table stakes, not a feature, and any bank pitching it otherwise is selling something.
What's the cost of waiting instead? McKinsey puts the odds at roughly 30% that AI reshapes global banking substantially as agentic systems take over core workflows, with as much as $170 billion in global banking profits at stake for institutions that don't adapt. Delay carries its own risk here, and the frameworks above suggest the tools to move responsibly already exist. What's left is practice: the unglamorous work of actually building to the standard, one bounded mandate and one audit trail at a time.
Sources
- Agentic AI in Banking Guide | Creatio
- How Agentic AI Will Reshape Payments in: IMF Notes Volume 2026 Issue 004 (2026)
- Agentic AI in Financial Services: A Research Roundup for 2026
- Agentic AI in Banking: The trend moving beyond pilots
- How Retail Banks Can Put AI Agents to Work
- multimodal.dev
- acropolium.com


