RCSA and Three Lines of Defense When AI Agents Execute Transactions
Accountability fractures when AI agents execute transactions instead of humans.

Agentic AI has stopped just recommending things. Now it moves money, modifies accounts, completes transfers on its own, and that shift breaks the accountability logic that RCSA and Three Lines of Defense were built around. It's breaking faster than most institutions have adjusted their governance to match. I want to walk through where the old framework still holds, where it strains, and what has to change at each line, because the tools we already have can probably handle this, if we're willing to rework how we use them.
Adoption isn't waiting around for governance to catch up, either. S&P Global found 54% of financial-services firms had deployed AI as of January 2025, up from 40% a year before. Separately, 52% of financial services firms say they're in active agentic AI adoption, yet only 15% of CFOs say they're actually ready to deploy one. Governance, traceability, human oversight: those are the sticking points they cite. That gap between deployment and readiness is where risk quietly piles up, and closing it at the framework level, specifically for agentic execution (payments, transfers, account modifications) is what this piece is about. AI-assisted analytics and recommendation engines raise a different set of governance questions, and I'm setting those aside here.
What RCSA and Three Lines of Defense actually require at each layer, before AI enters the picture
RCSA is meant to run on an ongoing basis, not as a once-a-year audit exercise: identify operational risks, check whether the controls actually work, document the mitigation. Ownership has always sat with the first line.
The three lines divide labor cleanly, at least on paper. First line, meaning the business units, owns and runs the process; they self-assess their own controls and carry the primary accountability when something goes sideways. Second line, operational risk and compliance, sets the RCSA methodology and challenges what the first line reports. They monitor how controls perform, but they don't own the risk itself. Third line, internal audit, tests independently whether the controls work as designed and reports what it finds up to the board.
That whole division rests on one quiet assumption: a human sits at every node, and that human can be found, interviewed, asked to explain their reasoning. Bottom-up RCSA leans on control monitoring by the first line and control testing mostly by the second. It's a clean split, but it only works because someone's actually there to account for it.
And this system was already showing cracks before AI showed up. An RMA survey found only 20% of respondents use modern tools like AI within RCSA itself, and plenty flagged weak integration across the three lines as a persistent gap. So agentic AI is stepping into a structure that was already straining.
What was this framework built to answer? Who authorized this, what was the control, did it hold. Fair enough. But what does it leave open once a non-human actor is the one taking the action? Who owns a risk when the thing executing the process is an AI agent running across a vendor stack, with no single person who "did" the transaction the way a loan officer or a wire clerk once did?
The regulatory void SR 26-2 created — and what it actually requires banks to do
SR 26-2, issued jointly by the Federal Reserve, OCC, and FDIC on April 17, 2026, replaces SR 11-7 as the governing standard for model risk management. Here's the paradox sitting at its center: it explicitly excludes generative AI and agentic AI from its scope. Footnote 3 calls them "novel and rapidly evolving," and on that basis, leaves them out. At the same time, the rule requires each institution's own risk management and governance practices to cover exactly that gap.
That carveout is not permission to relax. Examiners are already asking every bank, regardless of size, how it governs the AI systems this rule doesn't explicitly touch. The absence of a rule isn't the absence of an expectation, and banks that treat it that way are going to have an uncomfortable conversation with an examiner eventually.
Two pieces of SR 26-2 reach directly into agentic deployments even with the carveout in place. First: vendor model accountability. A bank can't offload risk just because it bought the model instead of building it. Oversight obligations follow the use case, not where the model came from, and a vendor's attestation that its system is safe doesn't satisfy the bank's own obligation. Second: materiality-based oversight replaces the old annual revalidation cycle with risk-based monitoring tied to how material the model actually is. Treat this as a paperwork update and you're missing what it demands operationally: continuous attention scaled to risk, not a box checked once a year.
International regulators have moved faster in places, and it's worth studying their approaches even if you're a U.S. institution not directly bound by them. Some international regulators have drawn genuinely concrete lines, explicitly requiring human participation or oversight for credit approval, account opening, and approval of deposits, withdrawals, or transfers. A Bank of England and FCA survey found 46% of respondent firms had only a partial understanding of the AI technologies they actually use, largely because those technologies come from third parties; meanwhile 84% said they had a named accountable person for their AI framework. Read those two numbers side by side and you start to see accountability structures running ahead of technical understanding in a lot of shops. The EU AI Act brings transparency obligations into force August 2, 2026, requiring that customers know when they're talking to AI, with the harder Annex III obligations around credit scoring and AML pushed to December 2027. DORA, in effect since January 17, 2025, requires EU financial entities to log, classify, and report ICT-related incidents, including ones tied to AI.
So where does that leave U.S. banks while SR 26-2's agentic guidance matures? Two horizontal frameworks are worth building from right now: the NIST AI RMF, with its Govern, Map, Measure, Manage structure, and ISO/IEC 42001, the first AI management system standard an external auditor can actually certify against, published December 2023. Neither replaces regulatory guidance, but both give you usable scaffolding today, which is more than nothing.
Where the Three Lines of Defense model breaks down when an agent executes a transaction
The first line's whole ownership logic assumes the risk owner exercises judgment and can explain that judgment afterward. An agent executing a payment doesn't do either of those things the way a person does. The employee nominally responsible for the process might have zero visibility into what the agent decided, or why, at the exact moment it acted. Accountability stops sitting with one person and scatters across model layers and vendor stacks instead. "Who approved this?" stops having a clean answer.
The second line runs into something related. Operational risk functions were built to challenge human self-assessments, people explaining their own reasoning back to a reviewer. They're rarely equipped to interrogate model behavior, tool-call logs, or an agent's reasoning chain. Control testing methods built around human processes can simply miss the failure modes specific to agentic systems: hallucination, adversarial prompting, one bad model output cascading into a hundred bad transactions before anyone notices.
Third line, audit, hits a wall that's almost mechanical. You can only test what got logged. If the agentic system doesn't produce structured, retrievable evidence of each decision and each action, there's nothing there for audit to examine. A lot of early agentic deployments simply don't have that infrastructure built in yet.
And then there's speed. Agentic systems execute faster than any manual monitoring cadence a human review cycle was ever designed for. So here's the real question: if a control can't be applied fast enough after the fact, does it belong after the fact at all? Mostly, the answer is no, and the control has to live inside the execution layer itself, not get bolted on afterward.
The lines themselves remain sound as a concept. What needs reworking is the ownership logic, the evidence requirements, and where the controls actually sit.
Reconfiguring the first line: who owns the risk when the agent owns the action
The fix isn't removing human ownership from the first line. It's redefining what that ownership covers.
Humans own the envelope the agent operates inside: the scope of what it's allowed to do, the thresholds and guardrails built into its configuration, the rules for when a human has to step in. The agent acts within that envelope, and accountability for the envelope itself stays with the first-line business owner, full stop.
What does that actually look like in an RCSA entry? A few things need to be documented formally, not just described in a narrative summary:
- Transaction type and value limits: the ceiling below which the agent can act on its own
- Counterparty or account restrictions: who the agent is and isn't allowed to transact with
- Escalation triggers: the specific conditions that force the agent to pause and hand off to a person
- Fallback behavior: what happens when the agent hits a step it can't complete, and whether it's designed to fail safe rather than fail open
Human-in-the-loop matters here, particularly for high-value or high-risk transactions. Emerging regulatory thinking on agentic payment systems is converging on human approval or supervisory intervention as a governance requirement once a transaction crosses into high-risk or high-value territory.
But here's a tension nobody solves easily: adding a human checkpoint can introduce liquidity risk if it delays a payment that needed to clear on time. Control design has to account for timing, not just who signs off at the end. And the kill switch itself can turn into a liability if it's built poorly, becoming a single point of failure, or worse, expanding the attack surface for anyone trying to manipulate the system.
At minimum, the first-line RCSA entry for an agentic process should capture the agent's defined authority, the approvals still required from a human, the controls built into the system itself, and the evidence showing those controls actually fired when they were supposed to.
This isn't some future problem for a lot of institutions, either. Cornerstone Advisors found in 2026 that 17% of credit unions have already invested in or deployed agentic AI, versus 7% of community banks. For a meaningful slice of the industry, reconfiguring the first line isn't on next year's roadmap; it's overdue now.
Reconfiguring the second line: how operational risk and compliance challenge what they cannot observe directly
The second line's job doesn't really change: set the standard, challenge what the first line reports, monitor how controls perform. What changes is what counts as evidence, and how the testing actually gets done.
Setting the standard means spelling out, in the RCSA methodology itself, what a valid entry for an AI-executed process has to include. Narrative self-assessment doesn't cut it anymore. What's required instead is documented agent authority, threshold logic, escalation rules, evidence the controls actually operated, plus vendor governance documentation, since SR 26-2 makes clear the oversight obligation follows the use case regardless of whose model it is. A vendor attestation alone won't satisfy that.
Challenging the first line's self-assessment gets harder when the "self" doing the assessing is partly a machine. One approach gaining traction, per BCG's 2026 analysis, is deploying dialogue-based AI assistants that help first-line risk owners refine their inputs and pressure-test their own assumptions before a formal challenge session even happens. The second line can turn the same tools around and use them to probe outputs rather than just accept them at face value.
Monitoring has to widen to cover failure modes that simply didn't exist in a human-run process: hallucination and model drift, where an output that was accurate at deployment quietly stops being accurate over time; adversarial prompting, where an agent working through a voice or text interface gets manipulated by a crafted input; cascading errors, where one bad decision inside a multi-step workflow multiplies across dozens or hundreds of transactions before anyone catches it.
GRC agents can help here too. They can run repeatable analysis across evidence, policies, controls, and remediation tasks, with humans supervising and approving what comes out the other end. For second-line teams without the headcount to manually review every AI-executed transaction, that's a real force multiplier, though it still depends on human judgment to mean anything.
One more thing belongs to the second line specifically: owning the AI inventory. A complete, current register of every agentic tool and agent running in production is the baseline requirement before any challenge or monitoring function means anything at all.
Reconfiguring the third line: what internal audit needs before it can test agentic controls
Audit's independence only matters if there's something concrete to independently examine. No structured, retrievable log of decisions and actions means audit has nothing to test against, no matter how sharp the auditors are.
At minimum, an agentic transaction execution system needs to produce a timestamped record of each action the agent took and what triggered it, the state of the agent's authority configuration at the moment it acted (what thresholds and restrictions were actually in force then), whether a human escalation fired and what happened as a result, and a record of any exception or anomaly the system flagged, along with how it got resolved.
From there, audit programs need test objectives built specifically for agentic systems, adapted from first principles rather than lifted from a checklist meant for human-run processes. Did the agent stay inside its documented authority on every transaction? That's scope adherence. Did the escalation triggers actually fire when conditions met the threshold that should have triggered them? That's whether the control operated as designed. Is the first-line RCSA documentation current, reflecting the agent's real configuration rather than some stale version from six months back? Does the vendor governance paperwork actually meet SR 26-2's standard that oversight follows the use case?
There's a design implication buried in here that's easy to underestimate: you can't retrofit an audit trail onto a system that's already live. Logging and traceability have to get built in before the agent goes into production, which means audit needs a seat at the table during design, not a summons after something's already gone wrong.
That's a big ask for most institutions right now. IBM data cited by Backbase in 2026 found only 8% of banks were developing generative AI in a truly strategic, enterprise-wide way as of late 2024. Most agentic deployments are still tactical: one team, one use case, one vendor contract at a time. Which means after-the-fact audit engagement is still the norm, not the exception. Reversing that isn't a technology upgrade so much as an organizational habit that has to change, and habits change slower than software ever does.
The RCSA entry itself: what the documentation for an AI-executed process must contain
A conventional RCSA entry follows a simple chain: process, inherent risk, control, control owner, residual risk, evidence the control actually operated. For an agentic process, every link in that chain has to stretch further than it used to.
Process can't just describe the transaction type anymore. It has to describe the agent's actual role inside it: which steps the agent executes on its own, and which still require a human to kick them off.
Inherent risk has to widen too. Alongside the conventional operational risks a bank has always tracked, the entry needs AI-specific categories: model error, adversarial manipulation, scope creep (the agent quietly doing more than it was configured to do), and vendor failure, since so much of this infrastructure sits outside the bank's direct control.
Control needs a split that didn't exist before: controls embedded directly in the agent's own configuration, like thresholds and escalation logic living inside the system itself, versus controls applied from outside, like a human reviewing a flagged transaction after the fact. Those are two different kinds of protection, and an RCSA entry that blurs them together isn't giving anyone, examiner or auditor, an honest picture of where the real safeguard actually sits.
The goal here is honesty about how much the thing being governed has changed, applied within a framework that's worked for decades, so the paperwork catches up before an examiner, or worse, an actual failure, forces the question for you.


