Est.
FeaturesLong read

What AI Governance Means When Agents Are Moving Real Money

Banks must redesign oversight for agents that execute transactions in milliseconds.

Editorial team · · 10 min read
Cover illustration for “What AI Governance Means When Agents Are Moving Real Money”
Features · September 30, 2026 · 10 min read · 2,161 words

I've spent a fair amount of time this year reading through bank pilot reports on agentic AI, and, one thing keeps jumping out. The gap between agents that recommend and agents that execute is huge, and, most of the industry hasn't caught up to what that gap actually demands. This piece is about that gap: what it takes to govern an agent that moves money on its own, when the old model of "flag it for a human to approve" no longer fits the speed the agent is working at.

How broadly agentic AI has already entered banking operations

99% of companies say they plan to put AI agents into production, yet only 11% actually have. I keep coming back to that number because it points to something specific. Data problems, governance gaps, and security worries are what's slowing things down, and those are much harder to fix than just moving faster.

That 11% figure undersells banking specifically, though. A 2025 MIT Technology Review survey of 250 banking executives found 70% of institutions are already running agentic AI in some form: 16% live, 52% in active pilots. Most of the industry is already past the point of asking whether this works.

So what are these agents actually doing all day?

They're running cross-border payment chains start to finish, from initiation through settlement monitoring. They're managing liquidity and prioritizing payments inside real-time gross settlement systems — essentially the cash management banks have always done, just compressed and with a lighter human hand on the wheel. They're handling AML compliance work, where agentic AI deployments have been associated with false positive reductions of 40% to 70% and SAR cycle time cuts of more than half, typically within 8 to 18 months of go-live. Meanwhile, they're drafting credit risk memos: one US bank saw a 20% to 60% productivity jump and a 30% improvement in credit turnaround time.

Agents are already sitting inside workflows where a mistake costs real money, and the governance question caught up to production systems a while back.

Where governance actually breaks down when agents execute transactions

Ask CFOs how ready they feel, and the answer gets uncomfortable fast. A July 2025 PYMNTS Intelligence report found only 15% felt ready to deploy agentic AI systems, and the reasons they gave were specific: traceability, human oversight, governance.

What's spooking the other 85%? Governance concerns are cited as a data challenge by 48% of organizations. The fear isn't hypothetical, either: 77% of AI incidents cause financial losses, and 55% cause reputational damage. That's measured after the fact, not guessed at beforehand.

Three failure modes show up once agents start executing instead of just suggesting.

Traceability collapse comes first. An agent that plans and executes across several steps on its own is genuinely hard to reconstruct after the fact, unless the audit infrastructure was built for exactly that purpose from day one. Teams that try to bolt it on later tend to keep chasing the gap without ever closing it.

Scope creep is second. Give an agent loosely defined boundaries, and it can end up executing transactions nobody actually intended, especially once several agents are handing work off to each other. The handoff is where things go blind, since nobody's watching the seam.

Third, there's a mismatch between how fast agents act and how fast humans review. Human review cycles were built for batch work, hours or days at a stretch, while agents act in milliseconds. That mismatch calls for a different oversight model entirely, one built around the pace agents actually operate at.

Regulators are starting to say this out loud too: once the population of agents grows past a certain point, checking every individual decision by hand simply stops being realistic. Banks need to design for that reality before it becomes a live problem, not after.

The governance infrastructure transaction-executing agents actually require

Banks already have a model for this, and it doesn't need reinventing, just extending. Three lines of defense, adapted for agents: the business unit that owns the agent is line one, responsible for spelling out exactly what it's allowed to do. Risk management is line two, watching the controls and keeping a model inventory that now includes the agents themselves, sitting right alongside the statistical models banks have tracked for years. Internal audit is line three, checking independently that agent behavior matches what was promised and that the audit trail actually holds up when someone pulls on it. Emerging regulatory guidance folds board-level oversight and AI risk committees into this same structure.

Underneath that sit four kinds of configurable controls.

Transaction type restrictions cover which categories of action the agent can take at all, while counterparty constraints cover who it's allowed to transact with. Velocity limits cap how many transactions, and how much total value, it can move in a given window. Channel restrictions cover which rails and interfaces it's allowed to touch.

Then there's the audit trail, and "complete" here means something specific: every input the agent received before acting, every decision step in the chain, every tool or system it called. All of it timestamped, with actor identity attached at each step, detailed enough that a regulator could reconstruct the whole causal chain without guessing at any point in the middle.

A global standard is taking shape here too, built around a set of sound practices that treat agentic AI as genuinely different from earlier generations of bank AI models. Build toward that now, because building toward something narrower just means redoing the work in eighteen months, and eighteen months is a long time to spend twice.

Worth saying plainly: deploying agents on the banking rails that already exist, instead of standing up parallel infrastructure next to them, makes all of this dramatically easier to audit. The transaction record already lives somewhere regulators know how to look.

Transaction thresholds and human-in-the-loop design for high-autonomy workflows

Regulatory thinking on oversight models converges on something fairly intuitive once you sit with it. Above a certain transaction value, you need human approval or dual authorization. Below it, once the number of agents outgrows what humans can individually review anyway, you shift to AI watching AI instead of a person watching each decision.

The Bank of Thailand went further in its 2025 AI risk-management policy and made the threshold binding, not advisory. Human participation is required whenever AI touches credit approval, account opening, or approving deposits, withdrawals, or transfers. It's the most specific regulatory line drawn on this so far, anywhere.

What does that look like running in production? Here's one live example. User input gets parsed into structured data, the system generates a transaction summary, and the user has to explicitly approve, decline, or edit before anything fires. That single gate, sitting right at the point of highest consequence, carries most of the weight. Confirming the action that actually moves money matters more than signing off on every step of reasoning that led there, and you keep a full audit trail without giving up the speed that made building the agent worthwhile in the first place.

A fraud detection case study out of a large, internationally active bank is probably the cleanest example I've seen of autonomous work and human oversight actually living together well. An agent monitors more than 80 million signals a day, across transactions, card activity, and digital banking, then proposes new fraud detection rules. The fraud analytics team reviews and approves every single one before it goes live, and nothing skips that gate.

The results: roughly three-quarters of the bank's card fraud rules are now written or updated by the agent, and fraud losses dropped 20% in the first half of the fiscal year. The system was built in-house in three months. The agent does pattern recognition at a scale no human team could match, and humans still hold the pen on anything that actually changes what the system does.

None of this is one-size-fits-all, either. Where you set the threshold depends on transaction type, counterparty risk, channel, and the agent's own track record for accuracy. A threshold that makes sense for a fraud-rule proposal is not the threshold you'd want for a wire transfer, and treating them the same is how banks get burned.

How network-level payment infrastructure is encoding agent governance into the rails themselves

Governance is also moving beyond individual banks, out to the payment networks themselves. Mastercard and Visa are building agent-specific controls straight into the rails.

Mastercard Agent Pay is one such initiative, designed to give agents a dedicated payment credential. Agents transact using a dedicated Mastercard credential with identity and limit controls built into the architecture. The governance is embedded at the architecture level, not bolted on after. The broader vision is rails redesigned for machine-to-machine commerce with identity and limit controls native to the network.

Visa has similarly moved to build agent-specific authorization controls into its network, extending the logic of existing human authorization frameworks to cover agents acting on behalf of cardholders.

These network-level controls set a floor, though, not a ceiling. Agentic tokens and per-session limits govern the payment credential, but they leave the agent's actual decision-making, and the bank's own audit obligations, mostly untouched. A bank still needs its own configurable controls, its own audit trails, its own human-in-the-loop thresholds sitting on top of whatever the network hands it. The network can stop a token from being misused; whether the agent's underlying reasoning was any good is a separate question, and it still falls to the bank to answer it.

A new problem is opening up as this scales, too. Banks now have to authenticate not just people, but the agents acting on their behalf. Fraud teams should expect criminals to start learning how to hijack or impersonate legitimate agents, the same way they learned to impersonate legitimate users years ago.

What SOC 2 certification and compliance-readiness actually signal about a deployment

SOC 2 certification is a real signal, and also a floor rather than a finish line. Passing it confirms the audit infrastructure exists and has been checked by someone outside the company.

It stops well short of the agent's actual decision logic, whether its transaction thresholds sit at the right level, and whether the human-in-the-loop design is any good. Those stay the institution's job, full stop, no exceptions.

So what does a real compliance-readiness check look like for an agent that moves money?

Can the bank reconstruct, in full, any transaction the agent executed, if a regulator calls tomorrow asking for it? Are transaction limits and authorization boundaries documented, version-controlled, and reviewed every time they change? Does the model inventory include the agents themselves, with a named business line owning each one? Is there a defined escalation path for the moment an agent hits a transaction it isn't authorized to make?

One question sits underneath all of it: does the deployment run under real institutional visibility into how it reasons, or does the reasoning stay hidden in a black box while only the outcomes get reported? Banks that can answer yes across that checklist, and that build on existing rails instead of standing up something parallel, walk into a regulatory exam from a genuinely stronger position.

What governance-ready agentic deployment looks like as a practical starting point

Banker skepticism about letting AI execute real transactions on its own makes sense to me. It deserves specifics in response, grounded in what the infrastructure actually does. Governance infrastructure is what actually earns trust here, more than a confident demo ever will.

A sequence that reduces risk while it builds confidence tends to look something like this. Start with a transaction type that's bounded, high in volume, and low in individual risk, something where the audit trail is easy to check and a mistake is recoverable rather than catastrophic. Put explicit human approval gates at the highest-consequence decision points before letting autonomy expand past them. Build the audit trail and the model inventory before scaling up; retrofitting that onto an agent population that's already live is a lot harder than designing it in from day one. Use the three-lines-of-defense structure to assign ownership clearly before the first production transaction ever fires.

That fraud detection case study is worth holding onto as a model: a complex, high-impact system built in three months, by a team that kept humans firmly in the approval loop for every call that mattered. Governance and speed work together when the architecture gets built carefully from the start; the tension only shows up later, when it wasn't.

Whether agentic AI belongs in banking isn't really the live question anymore, since the adoption numbers and the performance results already settled that. What's still open is whether governance gets built with the same seriousness as the capability itself. My guess is that the institutions treating governance as a continuous discipline built into the system, rather than a box checked by compliance after the fact, are the ones whose agents get trusted with more next year.

Sources

  1. symphonyai.com
  2. neurons-lab.com

More in Features