Est.

How to Mitigate Operational Risk in Banks Deploying AI

Banks must redesign controls for AI systems that act, not just respond.

Editorial team · · 10 min read · Updated
Cover illustration for “How to Mitigate Operational Risk in Banks Deploying AI”
Agentic Banking Execution · September 5, 2026 · 10 min read · 2,191 words

Agentic AI is not a faster chatbot or a smarter version of the automation banks have run for decades. It plans. It reasons across steps. It executes multi-step workflows across systems, and it needs little human involvement along the way. That single shift changes what a governance gap actually is. It stops being a configuration problem, something you patch with a setting or a policy update, and becomes a design problem built into the system from the start.

A study by Perwez, posted on SSRN, gives this a name: the Ambiguity Threshold Hypothesis. The finding is specific. Agentic AI performs reliably on low-ambiguity, rule-based tasks. But if you put the same system into a high-ambiguity, judgment-heavy environment, its outcomes turn unstable or outright adverse. Compliance, credit, and AML are exactly that kind of environment. They run on judgment calls, incomplete information, and regional variation, so a model can't apply clean rules to them mechanically.

Legacy controls were never built to watch something that acts. They were built to watch something that responds. A traditional banking tool waited for a prompt or followed a fixed rule set. An agentic system can chain tools together, call other agents, and act on information coming in from outside the bank, which multiplies the number of places something can go wrong. And here is where regulation hasn't caught up either: WitnessAI's 2026 guide points out that the OCC's revised Model Risk Management guidance explicitly leaves generative AI and agentic AI models out of its scope. The gap is regulatory as well as technical.

Think about what a bank's data loss prevention system actually watches for. It catches file transfers. It flags attachments. It has no way to inspect what leaves the building through a natural-language prompt typed into a chat window. A compliance officer pasting a customer's transaction history into an AI tool to get a faster answer isn't moving a file, so nothing fires. And inside the agents themselves, the accountability trail often doesn't exist in the first place: service accounts and API keys used by these systems are frequently not tied to any individual. A basic audit question, who did this and under whose authority, has no answer in the tools banks already have.

None of this means agentic AI is unsafe to deploy. It means the risk model is different, and pretending it's a faster version of what came before guarantees the governance built for the old model won't catch what the new one does.

Where operational risk enters: the four failure modes banks are encountering

Operational risk in agentic banking doesn't come from one weak point. It comes from four separate failure modes, and each one slips past a different control the bank already has in place.

Start with prompt injection. In a standard large language model, a successful injection produces a bad answer, something misleading that a human can catch and discard. Giving that same model the ability to call tools and take action means a successful injection can move funds, approve an exception, or expose account data. OWASP ranks prompt injection among the top risks for large language models generally, but the consequence changes entirely once the model can act on the world instead of just describing it.

Then there's the infrastructure connecting these agents to the bank's actual systems. The Model Context Protocol has become a standard layer, and it sits between AI agents and core banking systems. A security assessment by Equixly found that 43% of the MCP servers it examined were vulnerable to command injection. Some of the documented vulnerabilities involved hidden prompts that quietly exfiltrated sensitive data. One compromised server can reach several core systems at once, which turns a single point of failure into a multi-system event.

Shadow AI is the third mode, and it works by going around the bank's perimeter entirely rather than breaking through it. Most employees already paste work material into AI tools through personal accounts, so that traffic bypasses the bank's single sign-on and any CASB monitoring. Security teams lose visibility into what regulated data left the building and what happened to it next.

The fourth mode appears directly inside AML and compliance workflows. NHIMG's guidance on agentic systems notes that these tools typically sit between transaction monitoring, alert triage, case management, and reporting, and warns that they can produce inconsistent narratives, delay reviews, or trigger compliance failures when they lack local context or oversight.

Operational risk in agentic banking is not one problem but four compounding failure modes, each invisible to a different legacy control. They happen inside live interactions and live workflows, which is precisely the territory those tools were never designed to see. A patchwork fix, one tool bolted on for each failure mode, chases four different blind spots that each legacy control was never built to see. That's the case for rethinking the architecture rather than the checklist.

What the regulatory environment requires, and the compressed timeline banks face

Regulators have stopped treating this as a guidance conversation. They're enforcing now, and for some institutions, the clock already ran out on part of the runway.

In the United States, SR 26-2 replaced the long-standing SR 11-7 model risk management standard. Colorado's AI Act adds a state-level layer: it takes effect January 1, 2027, and requires disclosure for automated decision-making technology, including notifying consumers when they're subject to it. An amendment in May 2026 stripped out the impact-assessment requirement that was originally part of the law, replacing it with the current disclosure-based approach.

DORA became fully applicable January 17, 2025, and it adds operational resilience obligations that extend to AI-dependent workflows. The NIST AI Risk Management Framework reinforces a point that matters for timing: AI risk has to be managed across design, deployment, and ongoing monitoring, checked throughout rather than only at the moment a system goes live. NHIMG ties that framework directly to AML governance obligations specifically.

None of this is a matter of banks dragging their feet on compliance culture. The gap between how fast AI is being adopted and how ready the governance is sits in the structure of how these deployments happen. Vendor contracts make the structural problem worse. Banks signing multi-year AI agreements without pricing out what it costs to exit are locking in whatever governance architecture, or lack of one, the vendor happens to provide. An analysis from LLRX flags vendor viability and switching costs as strategic risks that most procurement processes simply don't account for.

These pieces together form a set of overlapping deadlines, some already active, some arriving within the next year, layered onto a governance model that was built for a slower, less autonomous generation of software. The next question: why can't banks just add the missing governance once a system is already running?

Why bolting governance on after deployment doesn't work

Governance added after an agentic system is already live has to work within whatever permissions, data access, and action boundaries that system already has. It can watch what the system does. It cannot go back and redesign what the system is capable of doing. That's the core limitation, and it explains why so many banks discover their control gaps the hard way.

NHIMG's analysis of AML deployments describes exactly this pattern: many security teams only discover an agentic failure after compliance staff have already relied on the agent's output inside a live investigation. The gap gets discovered through consequence, not through testing. By the time anyone notices, a decision has already been made on faulty information.

Deferred governance also tends to produce over-privileged agents as a default condition. If you don't build least-privilege design in from the start, an agent accumulates permissions to cover the broadest task it might ever conceivably need to perform, rather than the narrowest task it should be doing today. And that risk compounds with how deeply the agent is wired into core systems: a voice AI that can write back into core banking systems during a single call session carries a much larger failure surface than one that only reads data. That surface gets fixed at the moment the system is architected, before someone later decides to govern it.

The industry already has a case study in what deferred reckoning costs. The industry has a documented case where rebuilding came at significant cost: the Erica build showed that legacy banking services had to be rebuilt before a conversational AI system could call them properly. Institutions that put that work off don't avoid it. They inherit it later as an emergency instead of a planned project. You need an API-first microservices layer that connects existing databases to a conversational interface while it leaves the core system intact, and it only works if you choose it deliberately before deployment. Once an agent has already integrated directly with core systems, that layer can't be retrofitted in afterward.

The lesson here is less about caution and more about sequencing. A system can't be governed into having properties it was never built with. If governance has to be architectural, the next question is what that architecture actually looks like in practice, piece by piece.

What purpose-built governance architecture contains

Governance for agentic banking is a set of operational constraints built into the system before a single transaction runs through it. Each piece closes a specific gap identified above.

Least privilege and bounded tasks close the over-privileged-agent risk directly. NHIMG's AML guidance says you should limit agents to read-only access for evidence gathering unless a specific approval step exists, and you should keep suggestion separate from execution. This is the operational form of the 2026 International Scientific Exchange on AI Safety's principle of least privilege.

Traceable identity closes the accountability gap described earlier. Every action an agent takes needs to tie back to an accountable identity, because service accounts and API keys that aren't linked to individual oversight generate audit findings under segregation-of-duties rules.

Real-time logging closes the reconstruction gap regulators care about most. NHIMG's control pattern calls for logs detailed enough that a reviewer can reconstruct the agent's full reasoning path after the fact. That log is the exact audit trail that examiners under SR 26-2, the EU AI Act, and DORA will ask to see.

Human approval gates close the consequence-ownership gap. Filing a report, closing a case, reaching out to a customer, or anything that creates a regulatory obligation needs a human sign-off before it happens. Automation handles the execution work. People keep the authority over anything with real consequence attached to it.

AI inventory and risk tiering close the visibility gap at the enterprise level. The AI Framework for Community Banks, published by Verapath and hosted by the ABA in June 2026, includes an AI Adoption Stage Questionnaire and a Risk and Control Matrix built around 230 control objectives. The inventory records what each AI tool does, what data feeds into it, what decisions it touches, and who owns it, then applies governance controls sized to the actual impact of each tool. Only 16% of banks and credit unions currently have an enterprise-wide AI roadmap.

Compliance built into payment flows closes the reactive-investigation gap. Running fraud checks, sanctions screening, and risk rules directly inside real-time transaction flows, instead of after the fact, turns the control from something reactive into something preventive. The migration to ISO 20022 gives banks the structured data foundation that makes this possible at the scale of an entire payment network.

Adversarial testing before go-live closes the injection and exfiltration gaps described earlier. The OWASP Top 10 for Agentic Applications 2026, paired with the MITRE ATLAS adversarial threat matrix, gives red teams a documented catalog of tactics, including prompt injection, data exfiltration, and tool misuse, to test against before a system ever reaches production. Testing run after an incident is incident response, not testing.

How institutions are making this work in practice

These aren't theoretical prescriptions. Institutions of different sizes are already building governance in at the architecture stage, and the results are visible.

At Patelco Credit Union, you can see the governance posture stated outright. Kal Majmundar, Chief Technology and Transformation Officer, put it this way: "For us, the question isn't whether the risk is there; the question is how we build in the governance and expertise to manage those risks at the front end as we continue to innovate on behalf of our members." That framing matters because it treats governance as the condition that makes innovation possible, rather than a brake applied to it.

Plaid's Guaranteed Payments product shows what accountability looks like when it's built into the product itself rather than handled after a problem occurs. Launched May 19, 2026, the system approves ACH transfers instantly as it monitors a large base of accounts, and it takes on full financial liability if bank transfers run slow. If an approved payment fails, Plaid covers the loss. The accountability is a design decision made before the product ever processed a transaction.

Between a credit union stating its governance philosophy out loud and a payments company engineering liability directly into its product, the pattern is consistent. The institutions getting this right treat governance as part of the build.

Sources

  1. What are the risks of AI in banking? A 2026 guide
  2. Why do agentic AI systems create operational risk in banking when they touch AML workflows?
  3. The Rise of Agentic AI in Banking by Ahsan Perwez :: SSRN

More in Agentic Banking Execution