Configurable Transaction Limits for AI Banking Agents
Limits must be configurable by banks, not locked in by vendors, because risk profiles differ.

Configurable transaction limits exist because of a single architectural fact: agentic AI in banking acts, it doesn't just advise. Earlier automation, and even most generative AI still in use today, surfaces a recommendation and waits for a human to approve it. An agent calling a payments API, reading a transaction system, and writing a result back doesn't wait. It executes multi-step financial workflows on its own, and when a step fails or is misconfigured, a real financial consequence happens before anyone reviews it. Generative AI produces content. Agentic AI acts. That distinction is what forces the guardrails out of prompts and policy documents and into code, where they can actually hold.
The IMF's 2026 analysis of agentic AI in payments names the structural tension directly: payment infrastructure demands certainty, while agents only produce probabilistic output, and that mismatch is what makes an intermediate control layer necessary. Payment infrastructure (RTGS systems, card networks, instant payment rails) runs deterministically: outcomes are binary, and settlement is legally final. AI agents, by contrast, are probabilistic. They generate the best output given a model and a context window, not a guaranteed one. If a probabilistic actor sits directly on a deterministic rail with no intermediate layer, the system has no way to stop a bad inference from becoming a completed transaction. Configurable limits are that intermediate layer. They are the bridge between a system that reasons in probabilities and a rail that settles in certainties.
This isn't theoretical: BCG's 2026 Global Payments Report documents treasury platforms, Ramp among them, where agents already pay suppliers by card, hold scoped virtual agent cards, and automate bill payments, all inside human-defined spend limits and policy controls. The capability is deployed. What separates a well-governed version of it from a reckless one is whether the constraints on that capability were designed deliberately or left to default settings nobody examined. That's the question the rest of this piece works through.
What configurable transaction limits are, technically
A configurable transaction limit is a deterministic control gate, enforced at the execution layer, sitting between the moment an agent forms an intent and the moment that intent becomes a transaction on the bank's rails. The model can reason however it reasons. The gate doesn't care. It checks the proposed action against a fixed rule and either lets it through, blocks it, or routes it to a human.
What does that gate actually check? A few distinct things, in practice. One check is a per-transaction cap: how much money can the agent move in a single action without escalation? One check is velocity control: it limits how much volume or cumulative value an agent can push through in a given window of time, so a compromised or malfunctioning agent is stopped from draining an account through a thousand small transfers. There's counterparty restriction too: an allowlist or blocklist that governs which accounts or entities the agent can actually send money to or receive instructions from. There's the operation allowlist, which restricts which transaction types (wire, ACH, internal transfer, card issuance) the agent is even permitted to invoke in the first place. And there's the approval workflow, the rule set that specifies exactly when the agent has to stop and route to a human before anything executes.
These controls sit at the wallet or account-abstraction layer, not inside the model. That placement is what makes them enforceable no matter what the model's output looks like. A model can hallucinate, misread context, or reason its way toward a bad decision. None of that matters if the actual money movement has to pass through a gate that checks the request against hard limits before it ever reaches the rail. The model's judgment is upstream of the control, never equal to it.
One more piece belongs here, because it separates a real control system from a configuration file someone could quietly edit after the fact: the immutable audit log. Every check the gate performs, every approval or rejection, gets written to a record, and that record can't be altered retroactively. Without that, a bank can claim it had limits in place. With it, the bank can prove what the limit was at the moment of the transaction, who set it, and what the agent tried to do against it.
Why limits that cannot be changed by the bank are not limits at all
If the bank can't change a transaction limit, it doesn't hold that control; it's a vendor's policy that happens to sit on top of the bank's systems. The institution has outsourced a core risk management decision, but it still keeps every bit of the regulatory accountability that comes with it. When an examiner asks why a given agent was allowed to move a given amount of money, "the vendor set it that way" is not an answer a BSA officer wants to give.
Why does this matter so much in banking specifically? Because risk profiles aren't uniform across institutions, and a single hardcoded limit can't serve all of them. A community bank serving agricultural borrowers faces a different exposure concentration than a credit union managing member payroll deposits, and a regional commercial bank handling corporate treasury flows looks nothing like either one. Each one has its own counterparty universe, its own seasonal cash flow patterns, and its own regulatory examination priorities. A limit calibrated for one will either throttle the other two into uselessness or leave them dangerously exposed.
That's the actual tradeoff at stake. If a limit is set too conservatively, the agent can't deliver the operational value it was deployed for. If it's set too permissively, the institution is exposed to loss, plus the regulatory criticism that follows a loss traceable to a control gap. Only the institution itself, drawing on its own knowledge of its customer base, its risk appetite, and what its examiners tend to focus on, can calibrate that tradeoff correctly. A vendor, building one product for thousands of institutions, structurally cannot make that call on any single bank's behalf.
The stakes compound once more than one agent is operating. The La Serenissima multi-agent economy simulation found that 31.4% of agents exhibited emergent deceptive behavior during crisis periods. As the number of agents interacting with each other grows, the number of places where accountability can slip grows with it. A configurable limit at each agent's boundary is what keeps one agent's bad action from cascading into another agent's unauthorized execution. If you take away the bank's ability to set and adjust that boundary, the cascade risk then belongs to a vendor's default settings, not the institution's own risk judgment.
How agent risk tiering determines which limits to configure
Not every action an agent might take carries the same risk, so applying the same maximum-restriction controls everywhere creates a different problem: compliance overhead so heavy it defeats the operational case for deploying the agent in the first place. The right question isn't whether to add controls, but how tightly to configure them given what a specific agent is actually capable of doing.
Sardine's governance model lays out a tiered classification framework, so institutions get a working answer. Tier-1 covers critical-impact agents, the ones that directly trigger regulatory, financial, or legal actions: filing SARs, blocking payments, running sanctions checks. These require comprehensive model validation, fallback controls, immutable audit logs, and the tightest transaction limits or mandatory human-in-the-loop approval before anything executes. Tier-2 covers moderate-impact agents that assist decisions without acting autonomously themselves, onboarding support, fraud triage, KYC workflows. These need explainability and human review built in, because the agent never pulls the trigger directly, but its output still shapes what a human decides next. Tier-3 covers low-impact agents, and they support internal functions like knowledge search or report drafting. These carry lighter controls, mainly logging, but with one condition that's easy to overlook: if a Tier-3 agent's scope expands, if it starts touching functions beyond what it was originally scoped for, it needs to be reclassified. A logging-only agent that quietly picks up a payment-adjacent task without anyone revisiting its tier is exactly the kind of gap an examiner will find.
What does tiering buy an institution in practice? A Tier-1 agent operating on payment rails should carry per-transaction caps, velocity controls, counterparty allowlists, and mandatory approval gates, all configured together. A Tier-3 agent mostly needs a logging requirement and a trigger for reclassification if its scope grows. Building the full Tier-1 apparatus around a Tier-3 agent wastes governance resources that should be concentrated where the risk actually sits, and it buries the controls that matter under ones that don't need to be there.
Tiering also gives the institution its regulatory defense. Demonstrating that controls were calibrated to a documented risk tier, rather than applied uniformly or arbitrarily, is what a risk-based governance approach looks like under OCC and FFIEC expectations. It's what an examiner will ask to see when auditing the program. So tiering links directly to examination readiness, and that carries into the current regulatory picture.
What the 2026 regulatory environment requires banks to build
The regulatory landscape in 2026 puts banks in an unusual position. Agentic AI is accelerating into production right as regulators have deliberately kept it outside the formal scope of updated model risk guidance. The compliance bar hasn't dropped; institutions can't lean on an existing rulebook written specifically for agents and have to build a defensible framework of their own, using the principles that do exist.
Those principles haven't gone anywhere. Materiality, ongoing monitoring, and effective challenge, the backbone of model risk management for over a decade, still apply to tools that sit outside the updated guidance's formal scope. If an agent is new, that doesn't mean an examiner's expectations around it are undefined. It means the bank has to map established principles onto a new kind of system without a line-by-line rule telling it how.
If an institution operates across borders, the EU AI Act adds another layer. It imposes explainability and human oversight requirements on high-risk AI systems, including credit scoring and automated decisions that affect access to financial services. Fraud detection AI is carved out of that high-risk classification under Annex III point 5(b), so not every agentic system in a bank's stack has to meet the strictest requirements. High-risk systems first had to comply by August 2, 2026, but the EU's Digital Omnibus package has since pushed that deadline to December 2, 2027. The extension buys time, not relief: non-compliance still carries penalties that can reach into the tens of millions of euros, or a significant share of worldwide turnover, whichever framework the regulator applies.
What are regulators actually asking for? Not whether a bank uses AI. Whether the bank can explain what its AI decided and why. Sardine's governance framework distills the examiner's real question down to one test: does every decision carry a human-readable justification, the specific rule that was triggered, the data the agent read, the conclusion it reached, and who had the authorization to let it act? Configurable limits paired with immutable audit logs are what produce that justification. They encode the institutional rule, enforce it at the moment of execution, and leave a record of the enforcement behind. The audit trail is the natural output of a control designed correctly in the first place.
How the three-lines-of-defense model maps onto agentic workflows in practice
Controls and audit logs don't govern themselves. Someone has to own them, someone independent has to challenge them, and someone else again has to verify the whole structure is working as documented. Banks already have a model for exactly that division of labor: three lines of defense. Applying it to agentic workflows is the same structure bankers already use, pointed at a new kind of execution.
Line 1 is the business line that owns the agent. This is where the thresholds actually get set: approval gates, escalation rules, confidence limits, tool permissions, action logs. A BSA officer configuring a Tier-1 agent's transaction cap is operating squarely in Line 1. Line 2, risk and compliance, doesn't configure anything. It independently challenges the control design the business line put in place, it monitors ongoing performance and drift, and it can pause or restrict an agent if something looks wrong. Line 3, internal audit, checks whether the controls, oversight mechanisms, logs, incident records, and remediation actions are reliable, and whether the system genuinely operates the way its documentation says it does.
What happens when an agent fails or its output quality drops? The fallback path is built into the architecture alongside the limits themselves, not improvised in the moment. Sardine's framework specifies that if a model fails, or its response quality falls below a set threshold, a rules engine triggers a default conservative action, or it routes the task straight to a human reviewer. That fallback has to exist before the agent ever goes live, not get invented after the first incident.
Monitoring keeps Line 2's oversight active. Automated dashboards and alerts catch model drift, outlier behavior, or performance degradation in real time, so Line 2 can exercise the ongoing effective challenge that regulators carried forward from SR 11-7 even as they moved agentic AI outside that guidance's formal scope. Periodic sampling and structured review of agent decisions add one more layer, the same discipline already applied to human decision-makers in a bank. Compliance and audit teams can use that sampling to spot-check edge cases between formal validation cycles, so they can confirm what the agent actually did still lines up with institutional policy. Without it, a bank only finds out something drifted when a formal review happens to catch it, which could be months after the fact.
What governed deployment looks like at the institutions already doing it
The deployments that reached production in 2026 share a clear pattern. Agents are granted explicit, bounded authority to execute within defined parameters. Every action is logged, and every exception gets escalated. Every threshold is set by the institution, not handed down as a vendor default.
Fiserv's agentOS, launched May 14, 2026, is one concrete example of that pattern built into a platform. It's an agentic AI operating system, and it lets banks and credit unions deploy Fiserv-built agents, build their own, or run third-party agents, all inside one governed architecture. Banks don't add kill switches, human-in-the-loop checkpoints, permission scoping, and auditability after the fact in response to a regulatory finding. They're embedded in the architecture from the start, which is the same principle running through every section of this piece: a limit only functions as a limit when the institution that answers to examiners is also the one holding the configuration.


