How AI Execution Changes Operational Risk in Banks

I've spent enough time around bank risk committees to know when a room has stopped understanding what it's approving. That's what's happening with agentic AI right now. When AI moves from suggesting to doing, the risk doesn't shrink; it moves into the systems that configure, watch, and constrain the agent doing the clicking. That's the whole story here, and it needs a different kind of attention than most risk teams are currently giving it.
For years, the dominant AI use case in banking was pretty tame: flag the anomaly, surface the recommendation, draft the memo. A human still read it, still decided, still hit submit. Agentic AI breaks that pattern, because now it initiates the payment, executes the transfer, approves the credit line, or updates the account (start to finish, with nobody in the loop at the moment it actually happens). The human has moved upstream, where they set the rules the agent follows, or downstream, where they review what already happened after the fact.
Why does that matter more than it sounds like it should? Operational risk used to live at the point of human execution: the teller, the analyst, the ops person who double-checked a wire before sending it. Agentic execution removes that exact moment of judgment, and the risk relocates to wherever the agent got built, constrained, and monitored. Research on large language model agents backs this up too: agents show a higher rate of unsafe behavior than the base models underneath them, and because they chain tools and call other agents to get things done, one bad input can travel through a banking workflow before a single person gets a chance to look at it.
I'm writing this for risk officers and bank leaders who already know AI changes things. What I want to do here is get specific about where the new risk actually sits.
Where banks actually stand on agentic AI deployment right now
The adoption numbers run further ahead than most outside observers assume. A 2025 survey of banking executives by MIT Technology Review Insights found 70% of firms already using agentic AI in some form: 16% in production, 52% piloting. This stopped being a future-tense conversation a while ago.
More than 160 agentic AI use cases were announced across 50 of the world's largest banks in 2025 alone, and the early returns aren't just hype. Manual workloads dropped 30 to 50% in early deployments, and one US bank running agents on credit risk memos saw productivity gains between 20 and 60%, plus a 30% improvement in how fast credit decisions turned around.
Cornerstone Advisors' 2025 report adds a detail I keep coming back to: 28% of banks and 29% of credit unions planned to roll out generative AI tools for the first time in 2025. A lot of those institutions are walking straight into agentic territory without governance built to match, and the pressure to move isn't imaginary. McKinsey pegs $170 billion in global banking profits at risk for firms that don't adapt.
What should worry a risk officer isn't how fast adoption is moving. It's that operational risk frameworks were built assuming a human executes the transaction, and banks are now deploying into a world where an agent does, before the risk models underneath have caught up.
The mechanics of how agentic execution creates new failure modes
Three things turn ordinary risks into something with real teeth: speed, reach, and chaining. An agent doesn't pause for review the way a person naturally does, so errors compound before anyone notices. Agents often operate across multiple systems at once, frequently with more access than any single human operator would ever hold, and agents call other agents, other tools, other steps in sequence, so a misconfigured rule doesn't just produce one wrong answer; it moves through the whole chain.
Put those three together and you get a cascade problem. One error in an agentic payment workflow can trigger a transaction error, a reconciliation failure downstream, and a data exposure incident, all at once, all from the same root cause.
There's an example worth thinking through here, even outside banking. Early in 2025, a healthtech company disclosed a breach touching more than 483,000 patients, caused by a semi-autonomous AI agent that pushed confidential data into unsecured workflows while trying to make operations more efficient. The lesson is structural: any sector running agents with broad access and good intentions can end up in the exact same place.
Banking carries an extra layer that most other industries don't. When agents across different institutions run on similar models and similar data, a single macro signal can trigger a wave of nearly identical responses at once, amplifying past what the system was built to absorb. The market starts behaving less like millions of separate decision-makers and more like one coupled system moving in sync.
Underneath all of this sits a quieter problem: over-provisioning. Agents frequently get handed more system access than the task in front of them actually needs, and when something goes wrong, that extra access is exactly what turns a small mistake into a wide one.
For a risk officer, the shift is worth naming plainly. The old failure modes were human: fatigue, error, the occasional act of individual fraud. The new ones are architectural, and architectural problems need architectural controls, not another training module.
Where the risk actually migrates: the governance layer as the new load-bearing structure
In a human-execution model, governance catches the error after someone makes it. In an agentic model, governance decides whether the error happens at all, or whether it slips through undetected. That's the whole difference, really.
The old three-line-of-defense model still applies, but it needs adapting for agents specifically.
Line one is the controls built into the agent itself: transaction limits, thresholds that trigger escalation, confidence floors below which it won't act alone, logging at every step. Line two is the independent risk function, watching model performance, looking for patterns in what gets escalated, holding the authority to shut the agent down if something's off. Line three is audit and compliance, checking that lines one and two actually work and that the paper trail satisfies what regulators expect.
There's a fourth piece agentic systems add on top: documented fallback procedures and recovery targets, especially anywhere a workflow depends on a single large language model, one cloud provider, or one orchestration layer. Concentration risk belongs on the operational risk register now, not tucked into vendor management paperwork.
Configurability functions as the control itself. The ability to say what an agent can and can't do, under which conditions, with what approvals required before it acts, is where the risk actually lives now.
Audit trails have to change shape too. Logging the outcome isn't enough anymore; regulators want the inputs, the intermediate steps, the reasoning chain, since the question they'll ask isn't just what happened, but why the agent did what it did.
McKinsey's 2026 survey found only about a third of organizations report mature governance in place. Most banks, in other words, have deployed agentic capability on top of governance infrastructure that was never built to hold it.
Payments and transfer automation as the highest-stakes execution domain
Payments are where all of this moves fastest, and where a mistake shows up soonest. Mastercard rolled out Agent Pay in April 2025; Visa and PayPal both announced their own agentic payments capabilities that same year. PwC found agents can cut cycle times in purchase order processing and matching by as much as 80%. That's a real efficiency win, but it also strips out the human checkpoint that used to catch mistakes before they went through.
Fraud is the counterweight, and it's not small. Global fraud losses are estimated to climb from $23 billion in 2025 to $58.3 billion by 2030, a 153% jump, driven largely by fraud getting more sophisticated rather than just more frequent. AI-enabled fraud specifically (deepfakes, synthetic identities) surged 1,210% between January and December 2025, compared to a 195% rise in traditional fraud over the same stretch. Deepfake-powered fraud files went from roughly 500,000 in 2023 to roughly 8 million in 2025, and synthetic identity document fraud alone jumped 311% in North America in the first quarter of 2025.
That sets up an uncomfortable paradox worth sitting with for a second. The same agentic systems that can catch and block fraud at machine speed are also exposed to being spoofed or manipulated at that same speed, because speed doesn't pick a side.
There's a systemic angle too. When AI agents automate execution or liquidity timing across institutions simultaneously, that kind of concentrated automated activity raises real concerns about correlated behavior and system-wide amplification.
For risk officers working in payments, the controls need to be granular, real-time, and connected: per-transaction limits, counterparty validation, and escalation paths to a human that trigger on anomalies as they happen, not in a batch review three days later. Research found eight in ten highly automated firms name data security and privacy as their top concern, more than double the 39% reported by firms with less automation. Automation maturity and security worry seem to climb together, not apart.
What regulators are now requiring from institutions running AI on live transactions
The regulatory picture came together fast between 2025 and 2026. DORA became fully applicable on January 17, 2025, and its ICT risk management requirements now explicitly cover AI. The EU AI Act's prohibited-practices provisions kicked in on February 2, 2025; the high-risk obligations aimed at credit scoring AI are set to apply from August 2026.
In the US, updated interagency model risk management guidance updates the federal framework for governing AI models. Germany's financial regulator, BaFin, issued AI-specific ICT risk guidance in December 2025, framing AI squarely as an ICT risk management matter under DORA rather than something handled as an ethics or innovation side project.
Explainability has stopped being a nice-to-have. A Q1 2026 report from Wolters Kluwer found explainability and transparency were the single most cited regulatory concern among financial institutions surveyed, named by 28.4% of respondents. And here's the part that catches a lot of institutions off guard: the OCC's existing Model Risk Management guidance explicitly does not cover generative or agentic AI. Banks can't assume their current model risk frameworks already handle what they're deploying now.
Enforcement is already catching up to that gap. Regulators have made clear that AI-assisted lending decisions producing disparate outcomes will face scrutiny institutions can't deflect. That makes explainability a fair lending issue, not just a documentation exercise for the model risk team.
The Bank of Thailand's 2025 AI risk-management policy is unusually direct on this point: human participation is required whenever AI touches credit approval, account opening, or approval of deposits, withdrawals, or transfers. Other regulators are watching how that plays out, and the direction of travel across the EU AI Act, SR 26-2, and the Monetary Authority of Singapore's guidance all point toward the same demand: continuous monitoring, oversight that doesn't stop once the agent goes live, and documented behavioral controls across the agent's entire operating life, not just at launch.
I'll say this plainly, because I think it deserves to be said without softening. Institutions moving fast on agentic deployment without governance built alongside it should expect scrutiny, and it won't come from some hypothetical future audit. It'll come from a fair lending finding or a failed model risk management exam, sooner than most banks are currently planning for.
What risk officers and bank leaders must actually manage differently now
Governance functions as the primary risk control now, not a compliance overlay sitting on top of the "real" AI work. Risk officers need to own decisions about how an agent gets configured, not just review what it did after the fact.
Privilege scoping comes first. An agent should only have access to what its specific task requires, nothing more, and handing an agent broader access "just in case" is a risk decision that should get treated like one.
Escalation architecture is second. The exact conditions under which an agent stops and hands off to a person need to be designed on purpose, tested on a regular basis, and written down clearly enough that an examiner can follow the logic without a translator.
Third, audit trail standards have to expand. Logs need the inputs, the intermediate reasoning, and the outputs together, because regulators want the decision chain, not just the final number sitting at the end of it.
Fourth, concentration and dependency risk needs its own line item. When a workflow leans on one model provider, one cloud platform, one orchestration layer, that dependency is a real operational risk now, and it needs its own recovery plan attached to it.
Where an agent runs on a bank's existing rails, rather than requiring a whole new infrastructure build, the bank can apply the access controls, transaction limits, and audit systems it already has to agentic execution, instead of building a parallel governance structure from scratch. That's often the difference between governance that's actually usable and governance that exists mostly on paper, unread until the exam.
Established compliance frameworks and certification standards certify that the control environment around an AI system is auditable and holds up under scrutiny, not that the AI's judgment itself is sound. That's the lens risk officers should be using when they evaluate AI vendors: not "does this agent make good decisions" but "can I prove, to an examiner, exactly how and why it made them."
So where does that leave the skeptics? The question was never whether agentic execution introduces risk. It does, and so does human execution; it always has. The real question is whether the governance built around the agent can hold the risk where it now sits (in configuration, in controls, in the audit trail, rather than in a human checking every transaction by hand). Banks that build that governance alongside deployment, not after it, will be able to show examiners exactly how their controls work. Banks that treat the two as separate projects will find out how wide that gap is at the worst possible time to discover it.


