Est.

Banker Skepticism About AI Agents and How Institutions Respond

Bankers resisting AI agents have legitimate concerns regulators now require them to address.

Editorial team · · 13 min read
Cover illustration for “Banker Skepticism About AI Agents and How Institutions Respond”
Agentic Banking Execution · September 4, 2026 · 13 min read · 2,933 words

Bankers who hesitate on AI agents aren't behind the times. They're doing the job a bank hired them to do: protect deposits, manage risk, and refuse to move fast on something they can't yet audit. This piece argues that skepticism is a legitimate institutional reflex, deserving of real controls rather than more training slides, and that the banks pulling ahead are the ones designing controls around that skepticism instead of trying to talk people out of it.

At Backbase's Engage 2025 conference, CEO Jouk Pleiter asked the room a blunt question: who's tired of AI hype? A lot of hands went up. That's worth sitting with for a second, because the room wasn't full of technophobes. It was full of practitioners, the people actually running pilots and reading the vendor decks. Their fatigue traced back to a specific pattern: early AI projects were siloed. A chatbot bolted onto a support line here, a proof-of-concept in a lab there, none of it touching core operations. Slap the label "AI" on a demo, get the press release, move on. No wonder the room was tired.

There are two kinds of skepticism worth separating here, because they get treated as the same thing and they're not. One is healthy: grounded in real risk, real prior disappointment, and real gaps in governance that haven't been closed yet. The other is stalled: paralysis, not caution. Gartner's 2025 survey of CFOs and senior finance leaders found 16% of finance functions have no AI plans at all, and another 25% don't know how to get from planning to a first pilot. Both groups are real. But only one of them is doing something useful with its doubt.

What agentic AI actually does in a banking context, and why it raises the stakes for skepticism

Start with the difference between generative and agentic AI, because the whole argument for caution hinges on it. Generative AI writes things. It drafts a reply, summarizes an account, produces a report, then hands the output to a person who decides what happens next. Agentic AI acts. It initiates a payment, opens a case, routes a transfer, flags a compliance signal and follows through on it, without waiting for someone to approve each step along the way.

That's a meaningful distinction. A chatbot that tells you your account balance is one thing. An agent that moves money out of that account carries a different weight entirely, and the gap between those two is where most of the anxiety in this piece lives.

J.P. Morgan's 2026 Payments Outlook lays out two modes for how agents are starting to handle money movement. One is "transaction" mode: the agent acts on its own but only within narrow, predefined limits, think of it as a leash with a fixed length. The other is "orchestration" mode: the agent manages a multi-step purchasing or transfer workflow autonomously, coordinating several actions in sequence without a human checking in at each one.

Here's why that execution layer changes the stakes. A wrong answer from a chatbot is embarrassing. A wrong action from an agent is operational, meaning it costs real money, trips real compliance rules, and can't always be undone with an apology. Banking in 2026 has reached a point where trust in these systems is something institutions measure, not just something they hope for. And the leading indicator to watch isn't the big national banks. It's credit unions, which now outpace banks in conversational AI adoption according to research on agentic AI in credit unions. AI at credit unions is moving out of pilot programs and into the actual, risk-sensitive workflows: underwriting, fraud review, day-to-day account servicing.

That shift from pilot to production is exactly what turns three specific failure modes from theoretical to consequential.

The three failure modes that make banker caution structurally justified

CCG Catalyst's analysis names three ways serious AI deployments in banking go wrong, and each one maps to a risk a bank examiner already knows how to spot.

First: the model gives a confident, wrong answer. In a compliance or credit context, that's not a typo, it's a hallucination with legal weight behind it. Second: the agent acts without enough human review, meaning it crosses a line nobody explicitly drew, because the boundary between what it's authorized to do and what it actually did was never tight enough. Third: the vendor changes the underlying model without telling anyone, so the bank doesn't find out something shifted until a downstream process breaks and someone has to trace the failure back to its source.

Each of those failure modes lands on a different institutional nerve. Wrong answers hit model explainability and validation, which happens to be regulators' top concern right now. Unsanctioned action hits the control gap between authorized scope and actual behavior. Silent model changes hit vendor governance: who's keeping the audit trail when the thing being audited just changed underneath everyone?

What makes this worse than a normal software bug is how agents chain together. One agent calls a tool, which calls another agent, which triggers a workflow somewhere else. A single mispriced trade or duplicated payment can travel through several connected systems before the first human even notices something's off.

And the fraud backdrop makes the whole thing sharper. U.S. consumers reported losing $12.5 billion to fraud in 2024, a 25% jump from the year before, according to FTC data, with bank transfers and real-time payments among the channels criminals exploit most. Banks built their identity and fraud defenses around humans transacting with human judgment behind every click. Those defenses were never designed for an AI agent transacting with a customer's own credentials, moving at machine speed.

The structural problem underneath all of it: most credit unions and community banks running AI agents right now don't have a governance program built to match, with research indicating fewer than one in eight had one in place. The gap has less to do with which banks are careful and which are reckless, and more to do with deployment speed outrunning the speed at which institutions can build the oversight to keep up.

What regulators now require — and why the rules have changed faster than most institutions realize

The regulatory ground shifted fast, and a lot of institutions haven't caught up to how fast. DORA became fully applicable on January 17, 2025. The EU AI Act's prohibited practices took effect February 2, 2025. High-risk enforcement for credit-scoring AI under that same act starts August 2, 2026. Germany's BaFin, in guidance issued in December 2025, put AI squarely inside ICT risk management under DORA, treating it as a core supervisory concern rather than a separate innovation-lab category.

In the U.S., the interagency SR 26-2, the revised Model Risk Management Guidance, came out April 17, 2026, alongside OCC Bulletin 2026-13 and FDIC FIL-15-2026. SR 26-2 doesn't hand down one universal explainability rule. It actually places generative and agentic AI outside its formal scope, on the grounds that these tools are new and still changing, while telling institutions to apply existing risk management principles and figure out the right controls themselves. Banks have to reason through it on their own, working from principles rather than a checklist or safe harbor.

The CFPB's Winter 2025 Supervisory Highlights said it plainly: there's no advanced technology exception to federal consumer financial law. Doesn't matter how novel the tool is; fair lending law still applies. The OCC's May 2026 Semiannual Risk Perspective warned that AI is reshaping the cybersecurity threat landscape for banks in a way examiners are watching closely.

Explainability isn't a nice-to-have here. Whenever AI touches a credit decision covered by fair lending law, explaining that decision is a legal requirement, and the OCC, the Federal Reserve, and the CFPB have all said as much. A Q1 2026 Wolters Kluwer Banking Compliance AI Trend Report found 28.4% of institutions named explainability and transparency as their single most pressing regulatory worry, the top answer in the survey.

The clearest example of what human oversight actually looks like in a rule, rather than a principle, comes from Thailand. The Bank of Thailand's 2025 policy explicitly requires a human in the loop for credit approval, account opening, and approval of deposits, withdrawals, or transfers. U.S. regulators haven't written a single rulebook that specific yet. But examiners are already asking the underlying questions, about model documentation, data lineage, how specific an adverse-action notice needs to be, whether an explanation has actually been validated, and most institutions aren't ready to answer those questions for an agentic decision the way they can for a traditional underwriting model.

The adoption picture: where finance functions actually stand, and what separates movers from the stuck

The numbers complicate both the "AI is everywhere" narrative and the "banks are all stuck" narrative; each captures only part of the picture. Gartner's 2025 survey found 59% of finance functions are using AI, which is basically flat against 58% in 2024, after jumping from 37% in 2023. Growth stalled. It didn't reverse.

That stall isn't spread evenly. Two groups explain most of it: the 16% with no AI plans on the books, and the 25% stuck between planning and running an actual pilot. But here's the detail that cuts against the "disillusionment" story: 67% of the finance leaders already using AI say they're more optimistic than they were a year ago, per that same Gartner data. People who've actually deployed something are more bullish, not less.

Deployment is also further along than the skeptics in the room might assume. The 2025 EY-Parthenon GenAI in Banking survey found 77% of banks have launched or soft-launched a GenAI application, up from 61% who'd done so as of 2023. The gap that actually matters isn't adoption versus non-adoption. It's adoption versus governance readiness: McKinsey's 2026 survey found only about a third of organizations report having mature governance in place, even as deployment marches ahead of it.

What's actually holding people back, according to Gartner's data, has less to do with fear of AI and more to do with data literacy gaps, missing technical skills, and messy data. CB Insights' Q4 2025 enterprise survey backs this up: 65% of enterprises point to internal expertise gaps and 59% cite integration headaches as their top barriers, both ranking ahead of cost or capability worries.

So being stuck and being skeptical are not the same condition. The institutions actually moving aren't defined by a higher risk appetite. They're the ones that built governance infrastructure that makes risk something you can see and manage, instead of something you just hope doesn't happen.

How institutions that are moving forward are treating skepticism as a design constraint

Reframe skepticism as a spec sheet instead of a roadblock, and the whole problem gets more tractable. Skepticism, taken seriously, tells an institution exactly what controls need to exist before an agent earns the right to take a consequential action.

Three design habits show up again and again in the institutions handling this well. Bounded authority means an agent operates inside transaction limits, account scopes, and action types that are spelled out explicitly and can be checked later, not implied by a training set or a vendor default. Full audit trails mean every action an agent takes generates a record in real time that a compliance officer or examiner can pull up and question, not a log someone reconstructs after something's already gone wrong. Configurable escalation triggers mean the institution, not the vendor, decides the threshold at which an agent has to stop and hand a decision to a human.

CB Insights, in its Q4 2025 analysis, gave this pattern a name that borrows straight from an older banking concept: "Know Your Agent," a deliberate echo of Know Your Customer. The logic is the same. Identify who or what is acting, verify its authorization, and monitor its behavior over time, whether that's a person opening an account or an agent moving money on someone's behalf.

There are real examples of this working at scale. Some institutions have deployed multi-agent workflows that pre-screen credit applications and flag anomalies inside a defined scope, reducing manual first-pass review while keeping the audit trail intact; human attention shifts toward interpreting flagged cases and overseeing the system. Visa's real-time risk scoring for account-to-account payments scores a transaction in milliseconds using contextual data, then automatically approves, declines, or flags it, with fraud signals, sanctions checks, and risk rules built into that scoring step rather than bolted on afterward as a separate review.

Notice what these examples have in common: none of them tear out the existing banking infrastructure. PaymanAI, for instance, deploys AI agents to execute real bank transactions on a financial institution's existing rails rather than replacing them. They add agent capability on top of rails that already carry the compliance and audit architecture a bank already trusts. That's not a small detail; it's arguably the whole strategy. SOC 2 certification and a vendor's readiness to meet compliance standards are turning into baseline requirements for evaluation, not competitive edges. CB Insights found data privacy and security ranks as the top factor enterprises weigh when picking an AI vendor, ahead of cost and ahead of raw capability.

The execution layer is exactly where this distinction stops being theoretical. A chatbot suggesting a payment and an agent actually sending one are separated by exactly the kind of control described above, and that separation is what agentic AI in banking has to reckon with directly. That distinction — between suggestion and execution — is exactly the design constraint institutions building agentic systems have to reckon with directly.in practice: the same audit and governance layer that protects a bank's core systems now has to extend to whatever is making decisions on top of them. For a skeptical compliance officer, the signal that actually earns trust isn't a demo. It's a governance program an examiner can walk through.

What the adoption journey looks like in practice, from first pilot to operational trust

Where should a first pilot actually live? The Filene Research Institute's primer on agentic AI for credit unions points to contact centers, underwriting support, and fraud detection, high-volume, well-documented workflows where a human is already checking the work. That's not a coincidence. Those are the places where an institution can build a record of what the agent did and how a person reviewed it, before extending any real authority.

Starting narrow isn't caution for caution's sake. It's how an institution builds the evidence that convinces its own compliance team, and its own examiners, that the next step is safe.

A pattern shows up across practitioner accounts, and it tends to run in three stages. First, visibility: deploy the agent in a read-only or advisory role, where a human still executes every action, and let the audit trail build up before the agent gets any transaction authority at all. Second, bounded execution: hand over authority within narrow, explicitly set limits, with escalation thresholds already agreed on before the agent ever takes an action on its own. Third, expanding scope: widen what the agent can do only as the audit record piles up and the governance behind it holds, on the institution's own timeline, not whatever schedule the vendor is pitching.

Credit unions face a real staffing squeeze right now, rising digital expectations from members, more competition, and headcount that isn't growing to match. That pressure makes the case for agentic AI concrete. It also raises the cost of getting the governance wrong, because a short-staffed team doesn't have much slack to absorb a failure.

Voice and text interfaces tend to be the easiest entry point for front-line staff, mostly because they don't require new infrastructure and the interaction feels familiar even when the execution behind it is brand new. And the internal case for continuing writes itself once results show up: that 67% of finance AI users who feel more optimistic than they did a year ago are the evidence their still-skeptical colleagues are waiting for. They have audit trails. They have outcomes. They have something concrete to point to, instead of a vendor's promise.

One line from the practitioner conversation is worth holding onto: a healthy level of skepticism, but excited for the future. That balance is the target: measured confidence built on evidence. The institutions that get there treat governance as the thing that makes progress possible, not the thing standing in its way.

The governance infrastructure that turns institutional skepticism into durable trust

Here's the gap that hasn't closed: plenty of institutions are running AI agents right now without a governance program built to match. The hard part was never turning the system on. It's building the infrastructure that makes the deployment defensible a year from now, when an examiner asks questions nobody thought to prepare for.

Two components separate a real governance program from paperwork designed to look like one. Model documentation and lineage means every model or agent in production has a written purpose, a training basis on record, a validation history, and a change log, exactly what SR 26-2 scrutinizes even though it stops short of naming one explainability standard everyone must follow. Effective challenge means someone inside the bank, not the vendor, can question an agent's decision and override it; that internal expertise isn't a nice extra, it's the difference between a governance program and a governance slogan.

Put those two pieces together and skepticism stops being an obstacle to manage around. It becomes the design brief. The bankers raising their hands at that conference were asking, reasonably, to see the audit trail before they'd hand over the keys. That's the job.

Sources

  1. americanbanker.com
  2. ccgcatalyst.com

More in Agentic Banking Execution