Est.

Explainability Requirements for AI-Driven Banking Decisions

Regulators now demand banks prove they can explain every AI credit decision, or face real penalties.

Editorial team · · 12 min read
Cover illustration for “Explainability Requirements for AI-Driven Banking Decisions”
Compliance & Auditability · September 30, 2026 · 12 min read · 2,592 words

Explainability for AI-driven banking decisions has stopped being a nice-to-have and become something regulators actively enforce. That matters because a lot of institutions still treat it like a documentation chore, something a compliance team writes up after the model ships. Every AI-assisted decision a bank makes now, from a loan denial to a flagged transaction, carries live regulatory exposure if the bank can't explain it.

The path here wasn't one dramatic announcement. It built up over a series of supervisory moves, each one tightening the standard a little more. SR 11-7, issued by the Federal Reserve and OCC back in 2011, laid the groundwork: banks had to explain model outputs, validate performance, and document limitations in terms senior management and validators could actually follow.

In between, the CFPB sharpened the credit-specific side of the obligation. Circular 2023-03 made clear that ECOA's adverse action notice rules apply just as much to AI-based credit decisions as to any human underwriter's decision.

SR 26-2 itself, issued jointly by the Federal Reserve, FDIC, and OCC, replaced both SR 11-7 and SR 21-8. It requires banks to understand and manage their models well enough to actually manage the risk those models create, covering design, assumptions, data, methods, limitations, performance, and ongoing monitoring. It applies squarely to traditional statistical models and machine learning systems. Generative AI and agentic AI, notably, sit outside its scope for now, pending a separate framework still to come.

A Q1 2026 Wolters Kluwer survey of 148 financial institutions found that explainability and transparency ranked as the single most acute AI regulatory concern, ahead of bias and discrimination, ahead of data privacy, ahead of fair lending. Two international frameworks add deadline pressure on top of the domestic one. The EU AI Act classifies most credit scoring and underwriting AI as high-risk, though fraud detection AI is carved out of that category under Annex III, and banks and insurers using high-risk systems face substantial fines if they miss the August 2, 2026 compliance date. Colorado took its own path: SB 24-205 got repealed and replaced by SB 26-189, effective January 1, 2027, which labels many financial AI systems "high-risk" and mandates risk management practices, impact assessments, transparency measures, and bias mitigation, backed by real per-violation penalties.

What changes the temperature of all this is examiner behavior. Reuters reported that the OCC and Federal Reserve have started asking banks, during routine exams, to map out exactly where AI touches higher-risk functions: lending, know-your-customer checks, sanctions screening. Examiners are probing governance frameworks, guardrails, human oversight, third-party risk, contingency plans, and the whole operational picture. That's a shift from asking "are you ready for this?" to asking "prove it." And proving it means understanding exactly what the rules now demand a bank produce. CFPB's January 2025 Supervisory Highlights: Advanced Technologies Special Edition (Issue 38, Winter 2025) reaffirmed that using a black-box algorithm does not exempt an institution from providing specific explanations.

What SR 26-2 and ECOA together require banks to produce

Putting SR 26-2's model risk management requirements next to ECOA's adverse action obligations produces a two-part standard, one that most current AI deployments in banking don't actually satisfy.

The first part is about understanding the model itself. SR 26-2 requires that a bank understand model design, assumptions, data sources, methods, limitations, performance, and monitoring as something the institution can produce and update on an ongoing basis. The FFIEC IT Examination Handbook puts the risk in blunt terms: AI that lacks transparency, where nobody can say how inputs got translated into outputs, raises both compliance risk and operational risk. Credit unions face a parallel expectation, worked through NCUA's own risk-management framework for AI use.

The second part is about explaining the decision to the person it affected. Under ECOA and Regulation B, a bank that denies credit, closes an account, or changes the terms on an existing one has to give the customer a statement of the specific principal reasons behind that move. A generic reason doesn't cut it. Neither does "the model decided it." And it doesn't matter whether that model was built in-house, bought from a vendor, or hosted on someone else's machine learning platform: the bank still owns the obligation to explain it. Outsourcing the technology never outsources the responsibility. Fair lending reviews depend entirely on this kind of understanding, since regulators need to see why a model produces the credit access, pricing, and underwriting outcomes it does, and explainability is what makes bias identifiable in the first place.

A few other standards define what "explaining" is actually supposed to mean in practice. NIST, building on a definition from GAO, lays out four principles: an explanation should be supported, meaningful to whoever's receiving it, faithful to what the model actually did, and honest about the limits of what the model knows. FATF has flagged explainability specifically for AI used in anti-money-laundering and financial crime detection. The EU AI Act calls explainability a "fundamental obligation" for any system touching credit, insurance, or someone's livelihood, and the UK House of Lords AI Select Committee has said flatly that it's unacceptable for an AI system to materially affect someone's life without a satisfactory explanation attached.

The gap opens up here, though. SR 26-2 covers traditional statistical models and machine learning systems that aren't generative or agentic. It explicitly leaves generative AI and agentic AI out of scope, a hole sitting right where banking AI is moving fastest. That's not a small carve-out. It's a hole sitting right where banking AI is moving fastest, and it's still unresolved as of mid-2026.

Why agentic AI breaks the explainability frameworks banks already have

Agentic AI creates a kind of explainability failure that standard model risk management paperwork was never built to catch, and the SR 26-2 exclusion leaves banks without a clear compliance path exactly where they need one most. To understand why this works differently than a standard credit model, look at what these systems are actually doing under the hood.

Modern machine learning models can be extremely accurate while remaining internally illegible. Explainable AI, or XAI, refers to systems that can show why they generated a given output: what influenced it, how, and by how much. Black-box AI, by contrast, produces the output without giving anyone enough visibility to understand the approach, the limitations, the reliability, or the specific factors that drove that one decision. The Bank of England's Financial Policy Committee has pointed to the complexity of some AI models, combined with their ability to shift and adapt dynamically, as creating new challenges for predictability, explainability, and transparency.

Agentic systems don't just add to this problem, they change its shape. A predictive model's explainability question is fairly contained: why did the model produce this score or this prediction? An agentic system's explainability question is bigger: why did it choose this particular course of action, and what was the full sequence of steps it took to get there? When an agentic system denies a loan application or flags a transaction for review, the reasoning behind that call might run through a dozen intermediate steps. None of that appears in a standard output log.

That gap is already appearing in audits. The most common finding in 2026: a bank uses AI to generate adverse action notices, but the audit trail only captures the model's output score, not the specific factors that produced it. The bank can point to the fact that the AI made the call. It can't say why, at least not with the specificity ECOA demands.

The Apple Card case from 2019 does not prove discrimination happened, but it shows what the explainability gap actually cost the institution involved. New York's Department of Financial Services investigated after allegations appeared that the algorithm gave women lower credit limits than men with identical financial profiles. The DFS investigation ultimately found no unlawful discrimination. But that's almost beside the point. The real failure was structural: the institution couldn't demonstrate what was actually driving the disparity, couldn't tell a technical glitch apart from discriminatory model behavior, and couldn't go fix it with any kind of targeted governance because it didn't know what needed fixing. Accurate, fast, and scalable isn't worth much if nobody inside the bank can answer the basic question of why. SR 26-2's explicit exclusion of generative and agentic AI means banks deploying these systems must construct their own governance frameworks, though the guidance itself directs them to apply existing risk management and governance practices to any systems outside its scope, while still being fully exposed to ECOA, fair lending, and AML explainability obligations.

The fairness problem inside explainability that most banks are not tracking

What if a model produces an explanation for every decision, satisfies every standard fairness metric on its outputs, and still creates legal exposure? That's exactly the blind spot researchers have started mapping. A model can check every fairness box in its results while its actual reasoning process stays deeply unfair, something the research literature is now calling procedural bias.

Popoola and Sheppard, out of Montana State University, published this finding in the Journal of Artificial Intelligence Research, describing a gap sitting right at the intersection of algorithmic fairness and explainable AI. Their paper identifies three separate mechanisms that generate this kind of explanation inequity. One is representation-driven: the explanations reflect training data that underrepresents certain groups in the first place. Another is a mismatch problem: the post-hoc explainer, the tool bolted on after the fact to interpret the model, doesn't actually reflect what the model is really doing internally. The third is actionability-driven, and it's maybe the most practically important one: an explanation like "improve your debt-to-income ratio" is something some customers can act on and others structurally cannot, depending on their circumstances.

Giving every group the same explanation format, what researchers call attribution parity, is necessary but nowhere near sufficient. What actually matters is epistemic fairness: whether every group has an equal ability to understand the explanation and act on it.

The legal stakes here are concrete. AI systems that produce disparate impact on protected classes, without a documented and legally defensible explanation behind the disparity, create civil rights exposure under both ECOA and the Fair Housing Act. That doesn't lower the exposure. It just moves where it's coming from.

A lending model that penalizes applicants by zip code, correlated with race, is the textbook version of this problem. But the failure isn't that the model used a bad feature. The failure is that the institution had no way to detect it, because the model's reasoning stayed opaque the whole time. And there's a deeper structural wrinkle here too: per Popoola and Sheppard, explanation fairness is what's called an interventional quantity, something that can't be identified from observational data alone without causal assumptions baked in. In plain terms, that means standard post-hoc explainers, the tools banks are relying on right now, cannot reliably certify that their own explanations are fair. They can produce an explanation. They can't prove the explanation itself isn't part of the problem.

Where vendor relationships create a compliance blind spot banks routinely underestimate

A community bank that never writes a line of machine learning code is still fully on the hook for the explainability of whatever AI its lending, fraud, or compliance vendor is running on its behalf. Vendor opacity has become one of the leading sources of examiner findings, and it's easy to see why once you consider how the accountability actually flows.

The blind spot works like this: the bank doesn't build the model, but the vendor did, and outsourcing that technology never eliminates the bank's need for real risk management around it. Examiners know this now. They're asking pointed questions about how a bank uses its vendors, how client data gets safeguarded, whether "kill switches" actually exist to halt a system if it starts misbehaving. They're also digging into third-party risk and oversight more broadly: subcontractor exposure, contingency plans, what happens operationally if the vendor's system fails.

So what should a compliance officer actually be asking a vendor before the next contract renewal? A handful of practical XAI techniques that banks should require from vendors recur repeatedly in the frameworks built around this problem, per Abrigo's guidance. Can the vendor produce feature attribution or reason codes at the level of an individual decision, not just aggregate performance numbers across the whole portfolio? Does its audit trail capture the full sequence of steps an agentic system took, or only the final output? What confidence scores or uncertainty indicators does the system surface, and who inside the bank actually sees them? Does the vendor test its own explanations for fairness across protected classes, or does it only test the underlying model? And can the bank actually access the vendor's model validation and monitoring records, or is that locked behind a proprietary wall?

There are specific techniques worth requiring by name. Feature attribution shows which inputs drove a specific output and roughly how much weight each carried. Reason codes translate those model factors into human-readable language, the kind that can go straight into an adverse action notice. Confidence scores surface the model's own uncertainty, so an analyst knows when a decision needs human eyes before it goes out the door. For generative AI specifically, prompt and response traceability captures exactly what input produced exactly what output. For agentic AI, action logs record the full sequence of decisions and steps the agent took to get from a request to a completed task.

The idea that explainability has to come at the cost of accuracy is a myth worth retiring. Explainable models can hit high accuracy without carrying the opacity that deep learning systems typically bring. Banks don't have to choose between a model that works and a model they can explain. That tradeoff, where it does exist, is usually a design choice.

What the institutions currently deploying agentic banking AI are doing about explainability

Agentic AI deployments at major institutions and their core providers are arriving faster than most banks' governance frameworks can keep pace with. The institutions handling this well aren't treating explainability as a report generated after the system goes live. They're building it into the architecture before a single transaction runs through it.

An architectural requirement, baked in from the start, forces the system to produce a record of what it actually did, while a documentation exercise added after deployment can only describe what a system is supposed to do. A documentation exercise added after deployment can describe what a system is supposed to do. An architectural requirement, baked in from the start, forces the system to produce the record of what it actually did, at the level of individual reasoning steps, as a condition of operating at all. Given SR 26-2's requirement that banks explain model outputs, validate model performance, and document limitations, this kind of built-in transparency isn't optional for institutions running agentic systems on real banking decisions.

That's the shape of the problem banks are up against for the rest of this decade. The rules that exist, SR 26-2 and ECOA together, cover a lot of ground but explicitly stop short of the systems moving fastest into production. Closing that gap isn't going to come from a future circular alone. It's going to come from banks and the vendors they rely on deciding, ahead of the next exam, that a decision nobody can explain shouldn't be automated yet. PaymanAI, an agentic banking AI platform built for financial institutions, approaches this by pairing transaction execution with full audit trails as a design condition, not an afterthought.

Sources

  1. Explainability in AI: Why is it critical in banking?
  2. Fairness of Explanations in Artificial Intelligence (AI): A Unifying Framework, Axioms, and Future Direction toward Responsible AI
  3. Explainable AI vs. black-box AI in banking: What examiners expect and what to ask your vendor - Abrigo
  4. Explainable AI in Finance: What Regulators Actually Require in 2026
  5. Explainable AI in banking: Embedding explainability into AI decisions | DXC Technology
  6. EU AI Act for Financial Services: What Banks & Insurers Must Do
  7. The Fed - FRB: Supervisory Letter SR 26-2 on Revised Guidance on Model Risk Management -- April 17, 2026
  8. SR 26-2 Regulates Your Models, Not Your AI Agents: What Banks Need to Know - CIMCON Software

More in Compliance & Auditability