AnyLearn
All lessons
Businessadvanced

What a Regulator Actually Asks For

Regulated deployment is judged on evidence, not intent. This lesson covers the assurance vocabulary supervisors already use: three lines of defence, effective challenge, independent validation, and the model risk management tradition, including the 2026 shift from SR 11-7 to SR 26-2 and the gap it deliberately leaves.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 9

The question behind every question

A supervisor examining an AI system is not primarily asking whether it works. They are asking how you know, and whether anyone independent of the people who built it checked.

That reframing explains most of what follows. Controls that would satisfy an engineering team, we tested it and the numbers were good, do not satisfy a supervisor, because the assertion and the evidence come from the same source. What is being assessed is not the system alone but the organisation's ability to have caught a problem.

The corollary is uncomfortable and useful: a well-performing system with no evidence trail is in a worse regulatory position than a mediocre one that was properly validated, monitored and documented. Regulators cannot verify performance directly. They verify process, and infer.

This is not bureaucratic obtuseness. It is the only workable approach when the supervised population is large, the systems are opaque, and the supervisor arrives after the fact.

The rest of this cursus is about producing that evidence for systems whose behaviour is statistical, whose inputs are unbounded, and which change after deployment. All three properties make the established machinery fit awkwardly, and the awkwardness is the subject.

Full lesson text

All 9 steps on one page, for reading, reference, and search.

Show

1. The question behind every question

A supervisor examining an AI system is not primarily asking whether it works. They are asking how you know, and whether anyone independent of the people who built it checked.

That reframing explains most of what follows. Controls that would satisfy an engineering team, we tested it and the numbers were good, do not satisfy a supervisor, because the assertion and the evidence come from the same source. What is being assessed is not the system alone but the organisation's ability to have caught a problem.

The corollary is uncomfortable and useful: a well-performing system with no evidence trail is in a worse regulatory position than a mediocre one that was properly validated, monitored and documented. Regulators cannot verify performance directly. They verify process, and infer.

This is not bureaucratic obtuseness. It is the only workable approach when the supervised population is large, the systems are opaque, and the supervisor arrives after the fact.

The rest of this cursus is about producing that evidence for systems whose behaviour is statistical, whose inputs are unbounded, and which change after deployment. All three properties make the established machinery fit awkwardly, and the awkwardness is the subject.

2. Three lines of defence

The organising model in regulated industries is the three lines of defence, and knowing it is the fastest way to make an AI governance conversation legible to a risk function.

The first line owns and manages the risk: the business unit that builds or deploys the system and makes the decisions it informs. They are accountable for the outcome.

The second line oversees and challenges: risk management and compliance, independent of the first line, setting the framework, reviewing what the first line did, and having the standing to disagree.

The third line provides independent assurance: internal audit, independent of both, reporting to the board or audit committee, verifying that the first two are functioning.

The placement question from the AI governance role cursus resolves cleanly here. AI governance work is usually second line, and where an organisation puts it inside the first line, the structural weakness is not a matter of seniority or personality. It is that the same line is proposing and challenging.

For an organisation without this structure, the vocabulary is still worth adopting, because it names the thing that actually matters: someone other than the builder has to be able to say no, and be heard.

3. Effective challenge

The single most useful concept imported from model risk management is effective challenge, and it has a precise meaning that organisations routinely dilute.

Effective challenge is critical analysis by objective, informed parties who can identify limitations and assumptions and produce appropriate change. Three conditions have to hold together.

Incentive: the challenger must not be rewarded for the system succeeding. Someone whose bonus depends on the model shipping is not an objective party regardless of their intentions.

Competence: the challenger must understand the system well enough to identify what is wrong with it. Independence without competence produces a review that checks whether the form was completed.

Influence: the challenge must be capable of changing the outcome. A review that produces recommendations nobody must act on is documentation, not challenge.

The third condition is the one most often missing, and it is testable. Ask what happened the last time the challenge function objected. If the answer is that the objection was noted and the project proceeded unchanged, the function is decorative, and a supervisor asking the same question will reach the same conclusion.

Applied to AI, effective challenge is what turns a team's confidence that the model is fine into an assessment someone else can rely on.

4. The model risk tradition

Financial services has governed statistical models under supervisory expectation for over a decade, and that tradition is the most developed body of practice available for AI assurance.

SR 11-7, issued by the Federal Reserve and the OCC in April 2011, became the de facto standard well beyond US banking. Its structure: sound model development and implementation, effective challenge through independent validation, and governance with clear board accountability.

Its core requirement is independent validation. Someone other than the developers must evaluate conceptual soundness, meaning whether the approach is appropriate for the purpose at all; perform ongoing monitoring; and conduct outcomes analysis comparing predictions against what actually happened.

The vocabulary matters for anyone doing AI governance outside financial services too, because supervisors in other sectors increasingly borrow it, and because the problems it solves, opaque quantitative systems making consequential decisions, are the same problems.

The tradition also supplies a useful discipline: a model inventory, tiered by materiality, with validation depth proportionate to consequence. That is precisely the inventory and classification pattern the AI governance framework cursus builds, arrived at independently two decades earlier.

5. SR 26-2, and the gap it leaves

On 17 April 2026 the Federal Reserve, the FDIC and the OCC jointly issued SR 26-2, Revised Guidance on Model Risk Management, which supersedes and replaces SR 11-7.

The revision preserves the core principles, validation, monitoring, governance and effective challenge, while moving to a more explicitly risk-based and materiality-sensitive approach, so that validation depth scales with the consequence of the model rather than applying uniformly.

The part that matters most for anyone reading this: SR 26-2 excludes generative AI and agentic AI from its formal scope, on the basis that these technologies are novel and rapidly evolving, with regulators soliciting input on appropriate governance approaches for them.

Read that carefully, because it is the opposite of what most people assume. The framework an organisation would naturally reach for to govern a generative AI deployment explicitly declines to cover it.

So there is a genuine gap. Traditional predictive models sit inside a mature supervisory framework. Generative and agentic systems sit outside it, and the EU AI Act's high-risk obligations do not apply until December 2027. An organisation deploying a generative system in a regulated context today is not operating under a settled standard.

What that means practically is covered in the third lesson: you cannot cite a framework, so you have to reason and document.

6. Which regime reaches which system

The coverage map is patchier than most organisations assume, and knowing where your system sits determines what evidence you owe.

A traditional predictive model in a bank, a credit scorecard or a capital model, sits squarely inside model risk management, with a mature and specific set of expectations.

A high-risk AI system under the EU AI Act carries the Articles 8 to 15 requirements, but not until 2 December 2027 for standalone Annex III systems.

A generative or agentic system in a regulated firm falls outside SR 26-2's formal scope, and outside the AI Act's high-risk regime unless its use case puts it in Annex III.

What still applies to everything, regardless: sector conduct rules, non-discrimination law, data protection, and the general expectation that a firm manages its risks. Those never had a technology carve-out.

The gap in the middle is real, and it is where documented reasoning substitutes for a citable standard.

flowchart TD
A["Your AI system"] --> B["Traditional predictive model in a bank"]
A --> C["High-risk under EU AI Act Annex III"]
A --> D["Generative or agentic system"]
B --> E["Model risk management: mature expectations"]
C --> F["Articles 8-15, applying from 2 Dec 2027"]
D --> G["Outside SR 26-2 formal scope"]
G --> H["Gap: no citable standard"]
F --> I["Sector conduct rules, non-discrimination, data protection"]
E --> I
H --> I
I --> J["These apply regardless, with no technology carve-out"]

7. What the EU AI Act adds

For organisations in scope, the AI Act supplies obligations that the model risk tradition does not, and they are worth separating out because they are genuinely new rather than a restatement.

Human oversight as a design property. Article 14 requires high-risk systems to be built so they can be effectively overseen, including enabling the person to understand capacities and limitations, to remain aware of automation bias, to interpret output correctly, to disregard or override, and to stop the system. Model risk management assumes human review; the AI Act specifies what the system must do to make it possible.

Competence as a deployer obligation. Article 26 requires oversight to be assigned to natural persons with the necessary competence, training and authority. That is a requirement of result, and it names authority explicitly, which most internal frameworks leave implicit.

Fundamental rights rather than only prudential risk. Model risk management is concerned with loss to the institution. The AI Act's concern is harm to natural persons, which pulls in populations whose interests do not appear in a firm's risk register.

And logging as an obligation. Article 26 requires deployers to retain logs for at least six months where under their control.

An organisation with mature model risk management has most of the machinery and is typically missing the fundamental rights lens and the oversight design requirement.

8. Sector overlays

Beyond the general regimes, sector rules impose their own requirements, and they usually predate and outrank any AI-specific framework.

In financial services: conduct rules on treating customers fairly, requirements to explain adverse credit decisions, anti-discrimination provisions on lending, and operational resilience regimes covering critical systems and third-party dependencies.

In healthcare: medical device regulation where a system makes clinical claims, clinical governance and patient safety frameworks, and the professional accountability of the clinician who remains responsible for the decision.

In insurance: rules against unfair discrimination and the treatment of proxy variables, which is a well-developed body of practice that arrived at the proxy problem long before machine learning did.

In employment across most jurisdictions: non-discrimination law applying to selection decisions regardless of whether a human or a system made them.

Two consequences. First, a system can be fully compliant with the AI Act and still unlawful under sector rules, because the Act does not displace them. Second, the sector regulator is often the body that will actually examine you, and they will use their own vocabulary rather than the Act's.

So the practical starting point in a regulated firm is not the AI Act. It is asking the existing compliance function what their supervisor already expects, and fitting the AI work into that.

9. What evidence has to survive

A supervisor arrives after something went wrong, or on a routine examination, and asks about a decision made eighteen months ago. What has to exist.

That the system was assessed before deployment, by someone competent and independent of the builders, against a written standard, with the assessment recorded including what it found rather than only its conclusion.

That the limitations were known and documented, including where the system performs worse and on whom.

That oversight was designed rather than asserted, with the reviewer's competence, authority and available time evidenced rather than stated.

That the system was monitored, with the metrics, the thresholds set in advance, and what happened when a threshold was crossed.

That a specific decision can be reconstructed: what the inputs were, what the system output, who reviewed it, what they did.

And that when something went wrong it was recognised, escalated, and produced a change.

Notice that none of these are about the model's architecture. A supervisor is not going to review your embeddings. They are going to ask who checked, what they found, and what you did about it, which is why the next two lessons are about oversight and validation rather than about models.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why can a well-performing system with no evidence trail be in a worse regulatory position than a mediocre validated one?
    • Regulators penalise high-performing systems more heavily
    • Supervisors cannot verify performance directly, so they verify process and infer from it
    • Evidence trails are the only legally recognised form of testing
    • Performance claims are prohibited without certification
  2. Which condition of effective challenge is most often missing, and how can you test for it?
    • Competence; test by reviewing the challenger's qualifications
    • Independence; test by checking the reporting line
    • Influence; test by asking what happened the last time the challenge function objected
    • Documentation; test by counting the reviews produced
  3. What did SR 26-2 do, and what is notable about its scope?
    • It extended SR 11-7 to cover all AI systems including generative ones
    • It superseded SR 11-7 but explicitly excludes generative and agentic AI from formal scope
    • It rescinded model risk management requirements for smaller institutions
    • It applies only to EU-supervised institutions
  4. What does the EU AI Act add that the model risk management tradition does not?
    • Independent validation of models before deployment
    • A model inventory tiered by materiality
    • Ongoing monitoring and outcomes analysis
    • Human oversight as a design property of the system, and a fundamental rights rather than purely prudential lens
  5. Why is the AI Act not the right starting point in a regulated firm?
    • It has been superseded by sector-specific guidance
    • Sector rules predate it, are not displaced by it, and the sector supervisor is usually the body that will actually examine you
    • It applies only to providers, never to deployers
    • Its obligations are voluntary for supervised institutions

Related lessons

Law & Compliance
advanced

Proof: Disclosure, Presumptions, and the Complexity Rule

Strict liability is worthless if the claimant cannot prove a defect they never saw. Articles 9 and 10 answer that with a disclosure order, three presumptions of defectiveness, a presumption of causation, and a rule turning complexity into the claimant's ally. This lesson works through the cascade, the three-year and ten-year clocks, and what a defendant should be able to produce.

10 steps·~15 min
Law & Compliance
advanced

Who Pays, and For What Damage

The Directive builds a chain of liable operators so an injured person in the EU always has someone to sue. This lesson covers the manufacturer and component manufacturer, the importer and fulfilment service provider route, the distributor's one-month rule, online platforms, how a modification makes you a manufacturer, the heads of damage including data loss, and the exemptions.

10 steps·~15 min
Law & Compliance
advanced

Defectiveness: The Safety a Person Is Entitled to Expect

A product is defective when it lacks the safety a person is entitled to expect. Article 7 turns that into circumstances a court weighs, several written for software: the ability to learn after release, interconnection, cybersecurity requirements, and recalls. This lesson works through the list, the rule that a later improvement is not an admission, and why compliance is not a defence.

10 steps·~15 min
Law & Compliance
advanced

Software as a Product: What the New Liability Directive Changed

Directive (EU) 2024/2853 replaces the 1985 regime and settles a forty-year argument by naming software a product. This lesson covers the new definition and why delivery method is irrelevant, why information is not a product, how components and related services extend the net, where open source sits, and why liability cannot be disclaimed by contract.

10 steps·~15 min