Est.
FeaturesLong read

How AI Agents Execute B2B Payments on Existing Bank Rails

AI agents route B2B payments through existing bank rails based on transaction urgency and cost.

Editorial team · · 10 min read
Cover illustration for “How AI Agents Execute B2B Payments on Existing Bank Rails”
Features · September 30, 2026 · 10 min read · 2,193 words

When an AI agent pays a vendor invoice, it picks from rails banks already run, ACH, wire, RTP, FedNow, and increasingly stablecoin rails, then executes through the bank's own infrastructure. Understanding how that choice gets made, and what has to be true for it to be made safely, is the gap between a real deployment and a demo that never leaves the sandbox.

What rail selection actually looks like when an AI agent makes it

Think about the call an accounts payable clerk makes fifty times a day without even noticing: is this urgent, is it big, does the vendor need it today or will next week do? An agent has to make that same call, except explicitly, in code, every single time.

The IMF's April 2026 note on agentic AI and payments breaks the job into pieces: initiation, routing, compliance checks, watching the payment settle, handling whatever goes sideways along the way. That's the full chain, and rail selection is just one link in it.

What actually feeds the routing decision? Transaction size against each rail's ceiling. How fast the payment needs to land, based on the counterparty's terms and the payer's cash position. Cost relative to the amount moving. Whether there's any chance the payment needs to be pulled back. Whether the receiving bank is even on RTP or FedNow, because plenty still aren't. ACH batch cutoffs. Whether the vendor is domestic or overseas.

None of that fits into a static routing table. A legacy system applies the same rule no matter what's happening that day, while an agent weighs current conditions and adjusts on the fly, which is the real operational difference.

Large language models bring something rule engines never had: they can read. A contract clause about early-payment discounts. An invoice with the due date buried in fine print. An email thread where a vendor mentions a cash crunch. A static system throws that context away, but an agent can act on it.

Agents also work across a whole payables batch instead of one invoice at a time. Should ten routine vendor payments get bundled onto ACH while one urgent one splits off to RTP? That used to be a judgment call made through manual triage, and now it's a decision an agent can make across the whole batch at once.

None of this is hypothetical anymore. Santander and Mastercard ran a live AI-agent-executed payment inside a regulated banking environment in 2026, using pre-authorized permissions and tokenized credentials, entirely within existing banking controls. BBVA with Visa and Nordea with Mastercard are running similar programs. Three separate efforts, same direction: this is becoming a pattern in European banking, not a pilot somebody quietly shelves in Q4.

How agents execute inside a bank's existing infrastructure, not around it

The agent doesn't touch the core; it sits above it.

Orchestration protocols like MCP (Model Context Protocol) act as a standard gateway, letting a compliant agent talk to live banking functions without an engineer writing custom integration code for every core system it touches. Nymbus and Mambu already have production MCP implementations running in banking environments.

So what does "executing on existing rails" actually mean day to day? The agent picks the rail, sets the parameters, then fires the transaction through the bank's existing payment processor or core connector. The bank's ACH origination agreement doesn't change. Its RTP membership doesn't change. Its FedNow participation doesn't change. The agent is using the bank's existing credentials, not requesting new ones.

Compliance runs the same course. OFAC screening, transaction limits, counterparty checks, those still run through the systems already in place. The agent calls them the way a human operator would, just faster, and without needing a coffee break at 2pm.

Backbase's AI-native Banking OS, out in April 2026, is a decent snapshot of where this is heading: one operating layer sitting above the core, payments, and CRM, where agents, employees, and customers work inside the same execution environment. A number of financial institutions are already clients. Temenos made a similar move in May 2026, announcing embedded AI agents across its core banking products, with Microsoft Azure handling the governed, multi-step workflows underneath.

For a bank sizing this up, the real question is an integration question. The rails stay, the accounts stay, the compliance stack stays. What's new is the layer of judgment sitting on top of it.

Why most B2B payment platforms weren't built to support this execution model

Banks aren't dragging their feet on purpose. Most payment platforms were built for batch processing and a human at an approval queue clicking "approve" one item at a time. Nobody built them expecting a machine to make hundreds of routing calls a minute.

Three cracks show up fast once an agent meets one of these older stacks.

Fragmented data is the first. An agent's rail choice is only as good as what it can see, and if counterparty records live in one system, payment history in another, liquidity positions in a spreadsheet somewhere else, the agent is making decisions half blind.

Unclear decision authority is the second. Legacy workflows assume a person signs off at every threshold, but agents need something more exact: a written, machine-readable policy spelling out what they can execute on their own and what has to stop and wait for a human.

Integration overhead is the third. Wiring an agent into ACH origination, RTP, and FedNow all at once takes real API surface area, and a lot of mid-size banks just haven't built it yet.

There's a pattern that keeps repeating across enterprise AI rollouts generally: looks great in the demo, stalls in production, because the plumbing underneath was never built for continuous, autonomous execution. IBM's data from late 2024 found most banks were still in what it called "tactical mode" with generative AI: small experiments in isolated corners, not a real enterprise-wide push. This isn't a small-bank problem; it's everywhere.

The last decade of fintech spending went toward digitizing workflow: approvals, notifications, document handling. The next phase has to digitize the actual movement of money, or agents just make faster decisions on top of a system still moving funds on the old clock. AI-powered often describes a human-driven process with AI bolted on, while AI-native describes a process built around agent execution from day one. That distinction is really what decides whether a platform can support this or just gesture at it.

The governance architecture that makes autonomous payment execution safe

Once a payment clears RTP or FedNow, it's gone. No recall, no clawback, no phone call to reverse it. If fraud controls are only checking after the money moves, they're checking too late, and the whole governance model has to shift to before execution.

There's a speed problem here worth sitting with for a second. A compromised agent with payment access could fire off dozens of small transactions in the time it takes a traditional batch fraud review to finish one cycle, and detection windows that used to run in hours now need to work in seconds.

The old fraud models don't transfer cleanly, either. They were trained to catch unusual human behavior, an odd login time, a strange purchase pattern. When the actor making the transaction is an agent, those baselines don't apply, so authentication now has to answer two questions instead of one: is this agent who it claims to be, and does it actually have the authority the person behind it delegated to it?

McKinsey's 2025 report on agentic AI governance frames this well, describing a shift from watching after the fact to what it calls "active safety engineering": monitoring, behavioral limits, and audit signals built into the agent's workflow itself rather than bolted onto the outside.

In practice, for a payment-executing agent, that means a spend limit scoped to each agent, each rail, each time window. Cryptographic proof of identity for both the agent and whoever delegated authority to it. A circuit breaker that halts execution automatically when the pattern looks off. An escalation path written into policy in a format the system can actually read and act on. And no agent touching a primary corporate account directly; payment authority lives in scoped sub-accounts instead.

Backbase's Sentinel layer, part of that April 2026 Banking OS release, checks every actor's permissions against the bank's own policy before anything executes, and logs the action. That's what it looks like when a vendor builds this into the platform rather than leaving it as a suggestion in a slide deck.

For a community bank or credit union, this isn't an abstract risk exercise. An agent without a governance layer around it is an untracked actor moving money on behalf of members, which is a fiduciary problem before it's ever an operational one. Humans set the rules; the agent executes inside them and flags whatever falls outside. That's the division that lets a board actually sign off on this.

What auditability means when an AI agent executes the transaction

Ask a compliance officer what keeps them up at night about agentic payments, and it's rarely the technology itself. It's proving, after the fact, exactly what happened and why.

The regulatory stack a bank has to satisfy is layered: SR 11-7 model risk guidance, GLBA, PCI DSS, NYDFS Part 500, and for anyone with EU exposure, DORA and the EU AI Act on top. DORA has covered banks' ICT third-party risk since January 2025, and the EU AI Act's rules for high-risk financial AI systems, covering transparency, traceability, and human oversight, had a compliance deadline of August 2, 2026.

A Q1 2026 Wolters Kluwer Banking Compliance AI Trend Report found explainability and transparency were the top regulatory concerns banks raised, ahead of bias or other governance issues. Regulators don't just want the right answer; they want the reasoning that got there.

So what does an audit trail need to hold? Every action logged with a timestamp, the agent's identity, who delegated authority to it, and why it made the call it made. The rail choice documented with the actual parameters behind it, not just "ACH was selected." Every override or escalation logged with the human who stepped in and their reasoning. Logs kept in a form examiners can pull directly, without somebody building a custom extraction script at 11pm the night before an exam.

Research published in Frontiers in Artificial Intelligence points to multi-agent compliance setups as a stronger fit than one do-everything model: one agent reads transaction patterns, another interprets the applicable regulation, a third scores the risk. Splitting the reasoning up that way seems to hold up better under the kind of scrutiny compliance work demands.

The FSB's 2025 guidance backs this up from the supervisory side, calling for real-time monitoring of agent behavior, automated alerts on anything anomalous, and maintained logs of AI activity. Regulators are treating this as an expectation now, and banks rolling out agentic payment execution in 2026 are stepping into a space regulators are already watching closely.

What banks and credit unions need to evaluate before deploying a payment-executing agent

The first question isn't which vendor to buy. It's which payments are even good candidates for an agent to run, because not every payment carries the same risk.

Break it down by type and the picture sharpens fast. High-frequency, lower-value, recurring vendor payments on ACH are strong candidates: the parameters are predictable, the multi-day settlement window gives you room for error, and the volume is high enough that automation actually pays off. Time-sensitive, high-value payments on RTP or wire can work too, but they need tighter controls up front, since there's no undo button once they clear. Cross-border vendor payments sit outside this for now; domestic real-time rails don't reach there, so agents can help with FX timing and routing logic, but the money still has to travel through SWIFT or a correspondent network.

Before any of this goes live, a bank needs honest answers to a short list of infrastructure questions. Is counterparty data, payment history, and liquidity information sitting in one place an agent can actually read, or scattered across five systems and a shared drive? Are the ACH originator, the RTP interface, and the FedNow gateway exposed through APIs an orchestration layer can call? Are spend limits and escalation rules written in a format the system enforces, not just described in a PDF procedures manual nobody opens?

On governance, a few things shouldn't be negotiable. SOC 2 certification for anything touching live execution. Agent credentials scoped tightly rather than given broad account access. An audit trail built in from the start rather than bolted on later. An escalation path that's actually been tested, not just described in a document that sits in a shared folder.

Most institutions will start on the cautious end, where the agent recommends and a person approves, then expand the agent's authority as it builds a track record. That's how trust gets built when the money moving isn't yours.

Automation and accountability aren't really in tension here, provided the architecture underneath is built right. The agent executes, the institution governs, and every step in between leaves a trail somebody can actually follow, without needing to guess.

Sources

  1. elibrary.imf.org
  2. thefr.com
  3. banksandbankers.com
  4. paymentweek.com
  5. forbes.com
  6. frontiersin.org

More in Features