How AI Agents Handle Cash Management on Bank Rails
AI agents now manage real-time liquidity decisions inside payment systems built for human judgment.

A cash manager's job comes down to one tension, repeated thousands of times a day: hold enough money to cover payments as they land, but not so much that it sits idle earning nothing. Get either side wrong and there's a cost. This piece walks through how AI agents now handle that job directly on a bank's payment rails, and what actually keeps that automation honest.
Three decisions repeat all day: how big the buffer should be right now, which payments jump the queue when cash is tight, and when waiting to settle costs more than just borrowing to close the gap. These decisions happen inside real-time gross settlement systems, RTGS for short, where large interbank payments clear one at a time, in full, with no netting and no undo button. Send it wrong, and it stays wrong.
Miss on either side and the damage is real. Over-reserve and capital sits idle instead of earning yield or funding something useful. Under-reserve and a payment fails to settle, which trips reputational damage that outlasts the day it happened. This job has always needed a trained human at the screen: the decisions depend on each other, the clock never stops, and the trade-offs shift minute to minute in ways no fixed rulebook captures cleanly. That's the benchmark. Now here's what happens when a machine tries to clear it.
What the BIS found when it put a generative AI agent in that seat
The Bank for International Settlements ran a 2025 working paper testing something specific: could a generative AI agent handle intraday liquidity management inside a simulated wholesale payment environment? The model tested was general-purpose, the kind that wasn't built for banking at all.
The agent took on three jobs, the same three a human treasury operator handles: keep precautionary liquidity buffers where they need to be, prioritize urgent payments when the queue gets long, and weigh the cost of holding cash against the cost of letting something settle late.
Here's the part worth sitting with. The agent reproduced key cash-management behaviors without any specialized training for the task. The capability came out of general reasoning, not fine-tuning on payment data. A model that reasons well enough, pointed at this domain, starts behaving like it already knows the domain.
The BIS didn't call this a green light. The paper's language was careful: routine tasks could "potentially" be automated, and the researchers were explicit that safeguards and human oversight belong in any real deployment. Nowhere in the paper does anyone recommend an autonomous rollout.
So the open question turns mechanical fast. If a general-purpose agent can pull this off in a simulation, what does it look like running on live rails, with real money, inside a bank that answers to examiners? That's the rest of this piece.
How agentic AI differs from earlier bank automation, and why the distinction matters for execution
Older bank AI didn't do anything, strictly speaking. It predicted, it flagged, it surfaced a recommendation and waited for a person to click a button. Static models, rules engines, dashboards: all useful, none of them capable of acting on their own.
FinRegLab's 2025 research draws the line clearly. What makes a system agentic is planning, adaptation, and tool orchestration, together. The agent doesn't just guess an outcome; it picks a sequence of actions, calls the relevant APIs, watches what happens, and adjusts.
That's a different animal operationally. An agentic system doesn't hand someone a suggestion to act on later. It takes the action itself, inside whatever boundaries got configured ahead of time.
Three capabilities make execution on rails possible. Tool use lets the agent call payment APIs, check account balances, fire off a transfer directly. Planning lets it string together a multi-step workflow (check the buffer, scan the payment queue, route the payment, log what it did) without a human directing every step. Dynamic adaptation lets it change behavior as intraday conditions shift, instead of waiting for the next scheduled batch run.
The gap between "AI that helps" and "AI that executes" is the gap between a tool sitting next to the job and a system doing the job. Cash management on live rails needs the second one, and most banks still aren't there.
The rail infrastructure AI agents operate on and what each rail makes possible
Three domestic rails carry most of this traffic, and each does a different job. ACH handles high volume at low cost and is the backbone of everyday US payment flows, processing more than 35 billion transactions worth roughly $93 trillion in 2025. Wire transfer gives finality for large interbank settlement, the rail where the RTGS decisions actually live. RTP settles in real time, in seconds, and by 2025 has reached a broad share of US deposit accounts.
The rail map isn't fixed; real-time options continue to expand alongside existing infrastructure.
What does an AI agent actually see across all this? A queue of pending payments, each carrying a settlement deadline, a cost profile, a priority tag. That queue is the raw material every downstream decision runs on.
Picking the right rail for each payment is itself a task an agent can take on: route to ACH, Same Day ACH, RTP, or FedNow depending on urgency, cost, and what the settlement actually requires. Sophisticated platforms are building unified orchestration layers to do exactly this kind of routing automatically.
The agent sits above the existing rails, deciding how to use infrastructure that's already there. Core banking systems don't change. What changes is who, or what, is pulling the levers.
The execution sequence: what an AI agent actually does step by step during an intraday liquidity cycle
Break the cycle into steps and the mechanics get concrete fast.
Monitor. The agent reads account balances, incoming payment notifications, and queue position across every connected rail, continuously, not on a fixed interval.
Assess the buffer. It compares the current liquidity position against a buffer floor set in advance, and flags the position the moment it gets close to that threshold.
Prioritize payments. Every queued payment gets scored, by urgency, by counterparty, by settlement deadline, by the cost of letting it sit. That score sets the execution order.
Route to a rail. Based on urgency and cost, the agent picks ACH, wire, RTP, or FedNow for each payment and fires the API call to start it.
Reallocate if the buffer's at risk. Maybe it draws on a pre-approved liquidity facility. Maybe it shifts funds away from a lower-priority position. Either way, it stays inside treasury and risk limits set by humans beforehand; it doesn't invent new limits on the fly.
Handle exceptions. An unexpected drain, a failed settlement, a counterparty acting strangely: none of that gets resolved on its own. It gets flagged for a person to look at.
Log everything. Every action, every input that fed a decision, every API call, written to an audit record in real time as it happens.
For cross-border payments, the same logic extends further. The agent can optimize routing through correspondent banks or local partners, and trigger compliance checks before anything gets sent, the kind of multi-step orchestration that agentic AI research describes.
Step back and the pattern holds: the agent handles the complexity of execution, but every parameter governing its decisions, the buffer floor, the risk limits, the escalation triggers, got set by a person, in advance, on purpose.
How controls are built into the execution layer rather than bolted on afterward
Deployment doesn't have to be all-or-nothing, and it shouldn't be. A bank can start with the agent just recommending liquidity moves for a human to approve, then let it execute on its own below a set dollar threshold, then widen that scope as confidence in its behavior builds over time. That's the autonomy gradient, and it's a far more realistic rollout path than flipping a switch on day one.
Four properties need to live inside the execution layer itself, not sit off to the side as an afterthought. Accountability means a named executive owns the outcome of every AI-initiated action; the agent acts on behalf of a person, never apart from one. Transparency means the system can explain, in plain terms, what inputs and logic drove a given action. Auditability means every automated step generates a record that can't be altered after the fact: what happened, when, why, which agent did it, under whose authority. Continuous validation means the models get tested on an ongoing basis, not certified once at launch and left alone.
The permission-check model is worth naming directly: an authority layer that checks every actor's permissions against bank policy before anything executes, and logs the action regardless of whether it succeeds or fails.
Operating on existing rails means the bank's own authorization and settlement infrastructure stays the final gate. The agent routes, the agent initiates, but settlement finality still runs through the exact systems the bank already governs and already reports on to examiners.
The IMF's 2026 structural argument for how this should be layered is worth understanding. Let agentic AI operate in an upstream "intent and orchestration" layer, deciding what should happen and in what order, while keeping strict, rule-based controls in the authorization and settlement layers below it. That separation preserves accountability without giving up the efficiency the agent brings.
Baseline security and audit certification requirements for vendors in this space have risen alongside the stakes of what these systems touch.
The systemic risks regulators and researchers are watching that banks need to understand
Algorithmic herding is the risk that comes up first in almost every conversation about this. If a lot of institutions run similar models reacting to the same market signals, they can end up moving at the same moment, for the same reason. Researchers warn this could produce synchronized liquidity demand across the system, amplified swings that feed on themselves, congestion on the rails at exactly the wrong moment.
There's a regulatory gap sitting underneath all of this right now, and it's the part banks tend to underplay. SR 26-2, the interagency model risk guidance set to replace SR 11-7 in April 2026, explicitly puts generative and agentic AI outside its scope, calling the technology "novel and rapidly evolving." Banks deploying this today are operating ahead of settled guidance, not behind it. That's not a footnote, that's the actual state of play.
Readiness hasn't caught up with intent, either. Readiness surveys consistently find that a much smaller share of organizations consider themselves fully equipped to control and secure agentic AI systems than the share that plan to deploy them anyway. That gap between intent and readiness is the real story here, more than any single number.
Regulators have flagged AI-related model risk, cybersecurity exposure, and compliance risk as active supervisory concerns, not hypothetical future ones.
Cyber risk deserves its own line. Advanced models cut the time and cost it takes to find and exploit a system's weak points. Systemic-level research warns that correlated failures across institutions running similar AI systems could disrupt payments and financial intermediation broadly, not just at one bank.
None of this argues for waiting, though. The risks are real, but they're addressable through architecture: layered controls, clear human escalation paths, autonomy that expands conservatively instead of all at once. Regulatory clarity may lag deployment by years, and banks that wait for it aren't avoiding risk. They're just delaying the point where they start managing it well.
Where banks and credit unions are in adoption right now, and what the gap between leaders and laggards looks like
The declared numbers look impressive at first glance. As of 2026, 92% of global banks report active AI deployment in at least one core banking function. But declared deployment and scaled, governed, operational use are two different animals, and that 92% figure is doing a lot of quiet work to hide the gap between them.
Industry analysis names the actual bottleneck: fragmented data foundations, legacy systems that don't talk to each other cleanly, and internal resistance that throttles initiatives before they scale. A lot of institutions are stuck running isolated proofs of concept with governance that wouldn't survive an examiner's second question.
Industry analysis sharpens that picture further: only a fraction of banks worldwide are actively using AI in a way that produces real competitive advantage. The rest sit somewhere in pilot-stage fragmentation, technically "using AI" in the way that shows up on a survey but not in a way that moves outcomes.
The sentiment has shifted, though. The conversation inside institutions has broadly shifted: AI skepticism has faded, and the question has moved from whether to deploy this technology to how to get it to show up on the P&L.
For a community bank or credit union, the real cost of sitting this out isn't falling behind some abstract trend. It's ceding operational efficiency and member service capacity to institutions that have already moved past the pilot stage and are running production systems.
U.S. Bank's move is worth pointing to directly. It launched an AI-driven cash forecasting tool delivered through its digital banking platform, a client-facing product from a major institution with AI embedded into treasury operations.
Inside most community institutions, the workforce splits into people excited about this and people wary of it, and that split is probably healthy rather than a problem to smooth over. Working through that tension, instead of suppressing it, is part of how durable governance actually gets built.
What responsible deployment looks like in practice for a bank or credit union starting now
The real decision isn't build versus buy. That framing is too broad to be useful. The sharper question: which vendor's architecture gives an institution the controls and the audit trail it can defend in front of an examiner, without a fight?
Five things worth demanding from any agentic cash management system before signing anything. Configurable autonomy limits, set by the institution, not defaulted by the vendor. Immutable audit logs, timestamped, tied to a trigger and the decision inputs behind it. Human escalation paths, with clear rules for exactly when the agent stops and a person takes over. Rail-agnostic execution, working across existing ACH, wire, and RTP connections without ripping out core infrastructure to make room for it. Compliance-ready architecture, SOC 2 certified, with explainability good enough to satisfy ECOA and Regulation B adverse action requirements if the agent's decisions ever touch credit.
The autonomy gradient works as a rollout plan, not just a concept. Start with agent-assisted decisioning, where a human approves each move. Move to bounded autonomous execution, where the agent acts on its own below a defined threshold. Expand scope only once the audit record shows, concretely, that the system behaves the way it was configured to behave.
Examiners aren't going to ask whether an institution uses AI. That question is settled. They're going to ask whether the institution can show what the system did, when it did it, and under whose authority. Institutions that answer that cleanly are the ones that scale without a regulatory interruption stalling them out.
Backbase, serving more than 120 financial institutions through its AI-native Banking OS and Sentinel governance layer, and Temenos, with governed Azure AI agents built into core banking workflows, represent one category of infrastructure worth evaluating here, alongside purpose-built agentic tools designed specifically for banking execution.
The case for moving now rests on three things sitting side by side: the BIS finding that general-purpose reasoning already handles this job in simulation, U.S. Bank's live deployment proving it works in production, and Gartner's prediction that AI-on-AI governance will be its own defined market segment by 2030. The infrastructure exists. The governance frameworks exist. Waiting doesn't make any of this simpler; it just delays the point where the learning starts.


