BSA and AML Compliance When AI Agents Process Payments
AI agents that execute payments autonomously create blind spots in anti-money laundering oversight.

For years, AI in anti-money laundering work has played a supporting role. It surfaces alerts, flags anomalies in a transaction stream, drafts case summaries for an analyst to review. A human still reads the output, still decides, still hits send. That arrangement is changing fast, and few banks have adjusted their oversight to match.
Agentic AI closes the loop. It interprets what a customer or system wants, picks a payment action, and submits it to a settlement rail, often with little or no human check at the actual moment of execution. The IMF frames this as a structural shift rather than just another notch on the automation dial: transaction initiation moves from "explicitly human instructions" to what it calls "agent-mediated decision making." That phrase is doing a lot of work. It means the agent is making the decision itself.
Why does that distinction matter so much to a compliance officer? Because BSA obligations attach to the act of executing a transaction, not to the act of reviewing one. KYC and customer due diligence obligations still apply at onboarding and through ongoing monitoring, but there's a wrinkle now: who, or what, actually "initiated" a covered transaction is no longer a clean question.
The law hasn't caught up to the question either. An AI agent is neither. When an autonomous system causes a compliance failure, liability is in a gray zone at most institutions, unresolved by the very statute that's supposed to govern the transaction.
None of this is an argument against building agentic payment systems. It's an argument for building them with eyes open. The rest of this piece works through where the existing BSA/AML framework holds up under autonomous execution, where it cracks, and what a bank actually needs to do to keep its program intact and ready for an examiner who's going to ask hard, specific questions.
Financial crime compliance is already an enormous, expensive machine. Global spending on it exceeded $200 billion in 2023 alone, something like 3% of global GDP arxiv.org United Nations. U.S. institutions carry a meaningful chunk of that load themselves, spending somewhere between $35 billion and $40 billion a year just on AML operations finance.yahoo.com EU AI Act.
And the fraud side of the ledger isn't shrinking to compensate. Global fraud losses now run past $190 billion a year, and compliance teams report spending up to 42% of their budgets chasing down false positives that turn out to be nothing assistents.ai FinCEN Financial Trend Analysis, 2024. Nearly half of a compliance budget, in some shops, goes toward proving a negative.
Money is following the problem. The transaction monitoring software market was valued at roughly $19.98 billion in 2025 and is projected to climb to $41.99 billion by 2030, a 16.02% annual growth rate that signals a market still finding its footing symphonyai.com LexisNexis, 2024.
Because the alternative isn't a return to slower, safer manual review. Instant payment rails have already compressed the fraud detection window from days down to milliseconds, and human review at the actual point of execution is becoming close to impossible to sustain at that speed. The real choice facing institutions is between controlled AI deployment with full auditability, built deliberately, and ad hoc automation that quietly outruns governance until an examiner or a fraud loss forces the issue. The second path is the one that actually creates risk.
How agentic payment systems work across the three layers where BSA/AML obligations live
To make sense of where things go wrong, it helps to break an agentic payment down into stages, the way the IMF's research does. There are three layers, and each one carries a different flavor of BSA/AML obligation.
Layer one is intent formation and orchestration. This is where the agent interprets what a user or a system wants and plans out the sequence of steps to get there. Layer two is authorization and control: the agent picks a payment method, decides an amount, identifies a counterparty, and submits the whole package for approval. This is where CTR triggers and structuring risk actually live. Layer three is settlement, the point where the transaction finalizes on whatever rail is carrying it, ACH, wire, or real-time payments. Because settlement is often final and irreversible, this is where SAR timing pressure becomes acute.
McKinsey, in research cited by the IMF, described agentic systems as "digital factories," places where humans only step in for exceptions, a fine operating model for efficiency but not sufficient for BSA/AML LexisNexis, 2024. That's not sufficient for BSA/AML, which demands more than exception-handling at the human level. A compliance program built around "the human only looks when something looks weird" is a program that misses the transactions that don't look weird until you line up ten of them side by side https://ir.nasdaq.com/news-releases/news-release-details/nasdaq-verafin-announces-launch-its-agentic-ai-workforce.
The IMF names a tension directly: AI decision-making is probabilistic and adaptive, while payment infrastructure runs on deterministic rules. A model that gets the right answer most of the time isn't good enough here. "Usually correct" is not a standard BSA examiners recognize.
Mapping obligations onto the three layers clarifies where to focus. KYC and customer due diligence risk sits mostly at layer one: who is this agent acting for, and under what mandate? CTR and structuring risk sit at layer two: what's being authorized, in what amount, across what stretch of time? SAR filing duty and audit trail integrity live at layer three: what actually got settled, when, and can the bank reconstruct why?
One practical detail matters more than it might seem: the payment messaging standard a bank uses. ISO 20022's structured, standardized data fields improve AML screening accuracy at settlement and cut down false positives, surfacing risks that older, less structured message formats tend to hide Forrester, 2023.
Multi-agent setups complicate the picture further. A single payment might pass through an orchestrating agent, a separate compliance-checking agent, and then an execution agent, each one a handoff, and each handoff a place where the audit chain can quietly develop a gap.
Where CTR triggers, structuring rules, and SAR filing duties break down under autonomous execution
An agent operating across a session, or across a stretch of time, can execute several transactions without any single human ever deciding to do so. FinCEN's October 2025 SAR FAQs make a useful clarification here: transactions that are near the $10,000 line don't automatically signal structuring and don't, by themselves, trigger a SAR LexisNexis, 2024. Fair enough for a one-off. But an agent that repeatedly generates near-threshold payments in a pattern is a fundamentally different analytical problem than a single customer making a single judgment call LexisNexis, 2024.
Who's on the hook when that happens? The CTR obligation itself falls on the institution, not on whoever, or whatever, initiated the transaction. That reframes the real operational question: does the agent's execution record feed completely and automatically into the CTR generation workflow, or does it create a blind spot the institution doesn't know exists until an examiner finds it?
Structuring risk gets stranger under autonomous execution. An agent tuned to optimize for payment efficiency could, without any criminal intent baked into its design, start breaking large transfers into smaller pieces. The intent doesn't need to be criminal for the pattern to read as structuring to an examiner looking at the output. Institutions need to be able to show the agent's actual decision logic behind any transaction sequence that resembles structuring. A log of outputs isn't enough. Explainability has to exist at the authorization layer itself, before the pattern ever forms.
SAR filing carries its own set of pressures. FATF and U.S. regulators have held a consistent line: AI should enhance human oversight for high-stakes calls like filing a SAR, not replace it. Verafin's agentic sanctions analyst offers a useful real-world example of where the line currently sits. It can tell false positives apart from actual hits, and it gives a decision recommendation, flagging alerts as "Acknowledge" or "Review." Its 2026 expansion lets it autonomously close out false positives without a human recommendation step first, meaningful automation.
FinCEN's continuing activity guidance from October 2025 gives institutions some room here: banks aren't required to manually review after every single SAR filing, and can instead lean on risk-based internal policies to monitor ongoing suspicious activity LexisNexis, 2024. That flexibility comes with a catch, though. Those policies have to be "reasonably designed to detect and report ongoing suspicious activity LexisNexis, 2024." An agentic system has to demonstrably meet that bar. As Wolters Kluwer's Jo Brown put it, automation doesn't reduce accountability. If anything, as the AI gets more sophisticated, the oversight expectations placed on it go up, not down.
There's a related documentation wrinkle. FinCEN's October 2025 no-SAR FAQ confirms there's no requirement to write a memo explaining every decision not to file LexisNexis, 2024. But if an AI agent made or influenced that no-file decision, institutions ought to think hard about whether their existing risk-based documentation standard actually captures that adequately LexisNexis, 2024. A silent no-file decision made by a human analyst carries different institutional knowledge than one made by an agent nobody asked to explain itself.
And on the ACH side specifically, NACHA's 2026 rule requires all participants to document and implement fraud monitoring. Agentic ACH execution has to sit inside that documented monitoring framework symphonyai.com LexisNexis, 2024. It can't run alongside it as some separate, faster process that the fraud monitoring program wasn't built to see. CTR triggers involve a $10,000 cash transaction threshold.
The governance gap the 2026 model risk guidance leaves open
A sentence in the new guidance matters more than the update itself: generative AI and agentic AI models are described as "novel and rapidly evolving" and explicitly carved out as "not within the scope of this guidance." The agencies said a request for information on AI model risk management is coming "in the near future," which is regulator-speak for "not yet."
What that means in practice: agentic payment systems, the exact technology this piece has been describing, sit explicitly outside settled U.S. Bank model guidance right now, at the same moment banks are deploying them, leaves a defined, acknowledged gap rather than a fuzzy gray area.
The 2026 guidance applies most directly to banks holding over $30 billion in total assets OCC Bulletin 2026-13. Smaller, community banks work off a different baseline, but nobody gets an exemption from BSA obligations themselves OCC Bulletin 2026-13.
So what fills the space the 2026 guidance leaves open? A few things, partially. The U.S. Treasury's Financial Services AI Risk Management Framework lays out 230 control objectives spread across four functions: govern, map, measure, and manage U.S. Treasury Department. That's a concrete documentation structure banks can start using today, not a wish list. FinCEN's proposed risk-based program rule from June 2024 points in a similar direction, calling for institutions to build compliance frameworks around a formalized, ongoing risk assessment aligned to FinCEN's national AML/CFT priorities. It's not finalized, but the direction of travel is clear. Wolters Kluwer's guidance adds a useful principle for any bank buying rather than building: third-party AI tools carry the exact same accountability obligations as anything built in-house, meaning due diligence, ongoing monitoring, and service-level agreements covering accuracy, updates, security, and contingency planning.
One more data point to sit with. The OCC's Semiannual Risk Perspective found that AI is lowering the barrier to entry for threat actors, increasing the speed, scale, and sophistication of attacks against financial institutions. Oversight isn't optional just because the formal rulebook hasn't finished catching up. They won't, because the incentives and threats are moving faster than the rulebook can be finished. The practical move is to build governance now, using the Treasury framework, existing BSA program requirements, and the principles embedded in the 2026 guidance, even where it explicitly declines to cover the technology. On April 17, 2026, the OCC, Federal Reserve, and FDIC jointly issued OCC Bulletin 2026-13, replacing the 2011 model risk management guidance. The OCC's November 2025 update to community bank BSA/AML examination procedures aims to reduce unnecessary regulatory burden while maintaining strong oversight, and examiners now have explicit discretion to rely on a bank's independent testing if deemed satisfactory (LexisNexis, 2024).
The international regulatory pressure U.S. banks cannot ignore if they touch cross-border payments
Any bank moving payments across borders is dealing with more than U.S. rules, and the clock on some of this has already started running finance.yahoo.com EU AI Act. The EU AI Act's compliance deadline for high-risk systems under Annex III is December 2, 2027, and there's no grandfather clause protecting legacy systems already in production. AML transaction monitoring AI could fall under Annex III Point 5, covering essential services, or possibly Point 6, covering law enforcement contexts, depending on how the system functions and who's deploying it LexisNexis, 2024. That ambiguity itself is a planning problem.
The penalties attached to the EU AI Act aren't symbolic finance.yahoo.com. Prohibited AI practices can draw fines up to €35 million or 7% of global annual turnover EU AI Act. High-risk system violations top out at €15 million or 3% EU AI Act. Information breaches carry penalties up to €7.5 million or 1% EU AI Act. Those percentages are calculated against global turnover, not EU revenue, which changes the math considerably for a large multinational bank.
Colorado's AI Act, signed in 2024 and effective June 30, 2026, places requirements on developers and deployers of high-risk AI systems that materially affect financial services. That's a state moving ahead of federal clarity, and it likely won't be the only one.
FATF's late-2025 horizon-scanning work captures the dual nature of where things stand. AI opens up real opportunities for detection, but it's also arming criminals with new tools, including AI-generated deepfakes deployed against identity verification systems LexisNexis, 2024. FATF's position, consistent with everyone else quoted in this piece, is that AI should enhance human oversight in AML determinations, not replace it LexisNexis, 2024.
The EU has also built new supervisory infrastructure specifically for this finance.yahoo.com EU AI Act. Direct supervision of high-risk institutions under that authority begins in 2028.
And for banks eyeing stablecoin rails as a settlement option for agentic payments, the GENIUS Act removes a big piece of legal ambiguity: federally licensed payment stablecoin issuers will be explicitly treated as financial institutions for BSA purposes, effective January 18, 2027, possibly sooner.
When all of that is put together, the picture is clear: cross-border agentic payment execution was never just a BSA problem. It's simultaneously a BSA, EU AI Act, FATF, and FSB problem, running in parallel. A bank that solves for one and assumes the others will follow is building a governance architecture that satisfies exactly one regulator, and none of the others. The FSB's 2026 consultation on responsible AI adoption proposes sound practices for organization-wide AI governance and lifecycle management, explicitly asking whether practices address GenAI and agentic AI.
A practical framework for keeping BSA/AML programs intact when AI agents execute payments
Where does this leave a bank actually trying to deploy one of these systems responsibly? Start with mandate-based authorization. The IMF's research points to architectural separation of decision-making and execution as an emerging way to manage the risk: the agent's mandate, meaning what it's permitted to do, on whose behalf, and up to what limit, needs to be explicitly defined and technically enforced. Not assumed, not implied by training data, enforced. That means real, configurable controls sitting at the authorization layer: per-transaction limits, restrictions on which counterparties are eligible, and time-window aggregation caps designed specifically to catch the kind of inadvertent structuring pattern an efficiency-optimizing agent might create without meaning to.
Audit trails need to be built for this world, not retrofitted after the fact. Every autonomous action has to produce a timestamped, immutable record capturing the triggering input, the agent's decision logic at the intent layer, the authorization outcome, and the settlement confirmation. That record has to be readable by a human compliance officer. Nobody wants to reverse-engineer a model to figure out why a specific payment went out the door. Research out of Copenhagen Business School and the University of Copenhagen backs up that this is achievable: artifact-centric modeling, with clearly bounded agent roles and task-specific audit logging, shows compliance-by-design can be built into the architecture itself, as a core part of the design from the start.
Human oversight should be tiered, not applied uniformly. Routine, low-value payments that confirm an existing rule can run with the agent executing and a human reviewing after the fact. Payments that approach CTR thresholds, involve an unfamiliar counterparty, or trip any flagged pattern should pause for a human to authorize before anything settles. SAR filing decisions sit in a category of their own: a human makes that call, full stop, a point on which the IMF, FATF, and Wolters Kluwer all agree.
Any third-party agentic payment tool deserves the same scrutiny as something built in-house: ongoing monitoring, clear service-level commitments on accuracy, and a real plan for what happens when the system gets something wrong.
None of this is about slowing AI down for its own sake. It's about recognizing that the moment an agent moves from suggesting a payment to executing one, it steps fully into the BSA/AML framework that governs every other actor in the payment chain. The institutions that build for that now, deliberately, will be the ones still standing when the examiner finally asks to see the audit trail. The Treasury FS AI RMF's 230 control objectives across the govern, map, measure, and manage functions should be applied as the interim documentation standard (U.S. Treasury Department).
Sources
- How Agentic AI Will Reshape Payments in: IMF Notes Volume 2026 Issue 004 (2026)
- Finding the best AML transaction monitoring software 2026 - SymphonyAI
- Five developments every compliance leader needs to know | Wolters Kluwer
- AML Trends & Technology 2025: Turning Insights into Action
- arxiv.org
- How Agentic AI Will Reshape Payments


