AnyLearn
All lessons
Businessadvanced

Classifying a High-Risk AI System: Annex I, Annex III, and the Derogation

High-risk classification determines whether an organisation faces a substantial compliance programme or almost none. This lesson works through both routes: the Annex I product-safety route as narrowed in 2026, the eight Annex III use-case areas with the boundaries that get argued, and the Article 6(3) derogation, its conditions, and the assessment you must document to rely on it.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 10

The decision that sets the budget

Almost every question about the cost of AI Act compliance reduces to one determination: is this system high-risk?

On one side, a provider faces a risk management system, data governance requirements, technical documentation, logging capability, transparency toward deployers, human oversight design, accuracy and robustness requirements, a quality management system, conformity assessment, a declaration of conformity, CE marking, registration, and post-market monitoring.

On the other side, for a minimal-risk system, the Act imposes no product requirements at all.

There is no gradual slope between those positions. It is a threshold, and the entire apparatus switches on when it is crossed.

That makes classification the highest-leverage analysis in the whole regime, and it explains a pathology worth naming early: the incentive to classify downward is enormous, and the reasoning gets shaped by the answer people want. The defence against it is that the assessment must be documented and, in the derogation case, produced to authorities on request. A conclusion you would not want a sceptical reader to see is a conclusion to revisit.

Full lesson text

All 10 steps on one page, for reading, reference, and search.

Show

1. The decision that sets the budget

Almost every question about the cost of AI Act compliance reduces to one determination: is this system high-risk?

On one side, a provider faces a risk management system, data governance requirements, technical documentation, logging capability, transparency toward deployers, human oversight design, accuracy and robustness requirements, a quality management system, conformity assessment, a declaration of conformity, CE marking, registration, and post-market monitoring.

On the other side, for a minimal-risk system, the Act imposes no product requirements at all.

There is no gradual slope between those positions. It is a threshold, and the entire apparatus switches on when it is crossed.

That makes classification the highest-leverage analysis in the whole regime, and it explains a pathology worth naming early: the incentive to classify downward is enormous, and the reasoning gets shaped by the answer people want. The defence against it is that the assessment must be documented and, in the derogation case, produced to authorities on request. A conclusion you would not want a sceptical reader to see is a conclusion to revisit.

2. The Annex I route

The first route into the high-risk tier runs through existing EU product safety law, and it behaves quite differently from the second.

An AI system is high-risk under Article 6(1) where two conditions hold together. The system is intended to be used as a safety component of a product covered by the Union harmonisation legislation listed in Annex I, or is itself such a product. And that product is required to undergo a third-party conformity assessment under that legislation.

Annex I covers regimes including machinery, toys, lifts, equipment for potentially explosive atmospheres, radio equipment, pressure equipment, cableways, personal protective equipment, gas appliances, medical devices, in vitro diagnostics, and in a second part civil aviation, vehicles, marine equipment and rail.

Two 2026 changes matter. The Digital Omnibus narrowed the definition of safety component, and it introduced a carve-out for products covered by the Machinery Regulation, removing an overlap that had required manufacturers to satisfy two regimes for the same component.

The practical signal: if your product already carries a CE mark through a notified body, the Annex I route is live for you and the AI Act is layered onto machinery you already operate. If it does not, this route almost certainly does not apply.

3. The eight Annex III areas

The second route is the one ordinary commercial organisations meet. Article 6(2) makes systems in the areas listed in Annex III high-risk, and there are eight.

Biometrics: remote biometric identification, biometric categorisation according to sensitive or protected attributes, and emotion recognition, in each case where not prohibited outright.

Critical infrastructure: safety components in the management and operation of critical digital infrastructure, road traffic, and the supply of water, gas, heating and electricity.

Education and vocational training: admission and assignment, evaluating learning outcomes, assessing the appropriate level of education a person will receive, and monitoring prohibited behaviour during tests.

Employment and worker management: recruitment and selection including targeted job advertising and filtering applications, and decisions on promotion, termination, task allocation based on behaviour or traits, and monitoring and evaluating performance.

Access to essential private and public services: eligibility for public assistance benefits, creditworthiness evaluation and credit scoring, risk assessment and pricing in life and health insurance, and emergency call classification and dispatch.

Law enforcement, migration and border control, and administration of justice and democratic processes make up the remaining three.

For most businesses, only two of these are live: employment, and access to essential services.

4. Where the boundaries actually get argued

The list reads cleanly and applies messily. Four boundaries produce most of the real disputes.

Targeted job advertising is explicitly inside the employment category, which surprises marketing teams who assumed recruitment risk began at the application stage. A system deciding who sees a vacancy is deciding who can apply.

Task allocation based on behaviour or personal traits is inside; task allocation based on availability and skills is not. Workforce management tools sit on both sides of that line depending on what they optimise, and the distinction is a question about the model's features, not about the product category.

Creditworthiness evaluation is inside, with a carve-out for systems used to detect financial fraud. So the same institution can hold a high-risk scoring model and a non-high-risk fraud model, and the difference is what the output is used to decide.

Monitoring and evaluating performance is inside. A tool that summarises meetings is not evaluating performance. A tool that scores contributions is, whatever it is marketed as.

The pattern across all four: the category follows the decision the output feeds, not the technology or the vendor's product label. Classify the use, and reclassify when the use changes, which is exactly why the governance framework binds re-classification to new uses rather than to purchases.

5. The classification decision

Running both routes and the derogation in sequence.

Start by confirming the thing is an AI system at all, meaning it infers outputs rather than executing rules a person wrote. Then check the prohibited list, because a prohibited practice has no compliance path.

The two high-risk routes are independent, so both must be checked. Annex I asks whether the system is a safety component of a listed regulated product requiring third-party assessment. Annex III asks whether the use case falls in one of the eight areas.

An Annex I match ends the analysis: the derogation applies only to Annex III systems. An Annex III match opens the derogation question, which is the subject of the next steps.

Note the asymmetry. Annex I classification is largely mechanical, decided by which product legislation applies. Annex III classification is a judgement about what a system does to people, and it is where the analysis genuinely lives.

flowchart TD
A["Is it an AI system? Does it infer?"] --> B["No: outside the Act"]
A --> C["Check Article 5 prohibited practices"]
C --> D["Prohibited: no compliance path"]
C --> E["Annex I: safety component of regulated product?"]
E --> F["Yes: high-risk, no derogation available"]
C --> G["Annex III: one of the eight use-case areas?"]
G --> H["No: not high-risk by this route"]
G --> I["Yes: presumed high-risk"]
I --> J["Article 6(3) derogation conditions met?"]
J --> K["No: high-risk, full requirements"]
J --> L["Yes: document assessment, register anyway"]

6. The derogation and its four conditions

Article 6(3) is the pressure valve, and it is the most consequential and most misused provision in this part of the Act.

The general test: an Annex III system is not high-risk where it does not pose a significant risk of harm to the health, safety or fundamental rights of natural persons, including by not materially influencing the outcome of decision-making.

The Act then sets out conditions under which this applies. The system is intended to perform a narrow procedural task. Or it is intended to improve the result of a previously completed human activity. Or it is intended to detect decision-making patterns or deviations from prior decision-making patterns, and is not meant to replace or influence the previously completed human assessment without proper human review. Or it is intended to perform a preparatory task to an assessment relevant to the use cases listed.

And a hard limit sits over all four: a system that performs profiling of natural persons is always high-risk, regardless of which condition it might otherwise satisfy.

That profiling exclusion closes the route most organisations would want to use. Any system building a picture of an individual to evaluate personal aspects is out of the derogation, whatever its interface suggests about who decides.

7. Materially influencing the outcome

The phrase that decides most cases deserves working through, because organisations consistently read it too narrowly.

The intuitive reading is that a human making the final decision means the system did not materially influence it. That reading does not survive contact with how these systems are used.

A ranking system that determines which twenty of four hundred applications a recruiter reads has materially influenced the outcome, because the three hundred and eighty never considered were decided upon by the model. The human decided among the survivors.

A scoring system whose recommendation is followed in the overwhelming majority of cases has materially influenced the outcome, whatever the process document says, and the override rate is the evidence.

A system that flags applications for closer scrutiny has materially influenced the outcome if being flagged changes the treatment.

By contrast, a system that extracts structured fields from an uploaded CV so a human can read them consistently has not, provided it does not score, rank or filter.

The workable test: would the distribution of outcomes change if the system were removed and the same humans did the same job? If yes, it influences. Whether that influence is material is then a question about magnitude, and it is answerable with data you can actually collect.

8. Documenting a derogation assessment

Relying on the derogation is a positive act with obligations attached, not a quiet internal conclusion.

A provider that considers an Annex III system not to be high-risk must document that assessment before the system is placed on the market or put into service, remains subject to the registration obligation, and must provide the documentation to national competent authorities on request.

A defensible assessment covers six things.

The system and its intended purpose, stated precisely enough that a change of use would be visible as a change.

Which Annex III area it would otherwise fall into, named explicitly. An assessment that avoids naming the area reads as avoidance.

Which of the four conditions is relied on, and why the system meets it.

Why the system does not perform profiling of natural persons, addressed directly rather than by omission.

The evidence on material influence, ideally quantitative: what proportion of decisions follow the system's output, and what would change without it.

The conditions under which the assessment would need revisiting.

Write it for a reader who wants the opposite answer. That is the reader it may eventually get, and an assessment that only persuades someone already inclined to agree is not an assessment.

9. A worked classification

Three systems at one company, to show how differently the analysis runs.

SYSTEM A: CV parser
  Extracts name, dates, qualifications into structured fields.
  No score, no rank, no filter.
  Annex III area: employment (recruitment).
  Condition relied on: narrow procedural task.
  Profiling? No: no evaluation of personal aspects.
  Material influence? No: same applications reach the human.
  => Derogation available. Document and register.

SYSTEM B: applicant ranker
  Scores applications for fit; recruiter reviews top 20 of 400.
  Annex III area: employment (filtering applications).
  Profiling? Yes: evaluates personal aspects to predict fit.
  => HIGH-RISK. Derogation closed by the profiling limit.

SYSTEM C: interview scheduler
  Optimises slots against calendar availability.
  No inference about persons; executes constraints.
  => Not an AI system for these purposes. Record and exit.

System B is the instructive one. Every instinct says a human decides, because a recruiter reviews the shortlist. But the model chose the shortlist, and it profiles to do so. Two independent grounds put it in the high-risk tier, and no amount of process design around the human step moves it out.

10. Getting classification wrong in both directions

Under-classification is the failure everyone anticipates. Over-classification is the one that quietly wastes the budget.

Under-classifying carries the obvious exposure: operating a high-risk system without the requirements, with an assessment that will not survive scrutiny. It has a characteristic signature, an assessment concluding what the business wanted, written after the deployment decision rather than before.

Over-classifying is more common than expected. A team that treats every AI system as high-risk builds a quality management system it does not need, produces Annex IV documentation for a transcription tool, and exhausts the organisation's tolerance for compliance work before reaching the system that actually needed it. It also makes the framework less credible: when everything is high-risk, the classification carries no information.

The discipline that avoids both is the same. Classify against the text, area by area and condition by condition. Record the reasoning. Have someone other than the system's sponsor review the conclusion. And revisit when the use changes, because the classification attaches to the use rather than to the software.

With the classification settled, the next lesson covers what a high-risk determination actually requires you to build.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Which condition must hold, alongside being a safety component, for the Annex I route to make a system high-risk?
    • The product must be sold in more than one Member State
    • The product must be required to undergo a third-party conformity assessment under the listed legislation
    • The provider must be established in the EU
    • The system must also appear in Annex III
  2. Which limit closes the Article 6(3) derogation regardless of which condition a system might otherwise satisfy?
    • Systems deployed by organisations over 750 employees
    • Systems that process special category data
    • Systems that perform profiling of natural persons are always high-risk
    • Systems placed on the market after 2 December 2027
  3. A ranking system determines which 20 of 400 applications a recruiter reads. Has it materially influenced the outcome?
    • Yes, because the 380 never considered were effectively decided upon by the model
    • No, because a human makes the final hiring decision
    • No, because ranking is a preparatory task
    • Only if the recruiter follows the ranking without exception
  4. Under Annex III, which kind of task allocation falls inside the employment category?
    • Allocation based on availability and skills
    • Allocation based on behaviour or personal traits
    • Any allocation performed by software
    • Allocation in organisations with more than 250 employees
  5. What is the practical cost of over-classifying systems as high-risk?
    • Penalties for incorrect registration in the EU database
    • Loss of the right to rely on the derogation in future
    • Mandatory notified body involvement for all systems
    • Wasted compliance effort and a framework whose classifications stop carrying information

Related lessons

Law & Compliance
advanced

Proof: Disclosure, Presumptions, and the Complexity Rule

Strict liability is worthless if the claimant cannot prove a defect they never saw. Articles 9 and 10 answer that with a disclosure order, three presumptions of defectiveness, a presumption of causation, and a rule turning complexity into the claimant's ally. This lesson works through the cascade, the three-year and ten-year clocks, and what a defendant should be able to produce.

10 steps·~15 min
Law & Compliance
advanced

Who Pays, and For What Damage

The Directive builds a chain of liable operators so an injured person in the EU always has someone to sue. This lesson covers the manufacturer and component manufacturer, the importer and fulfilment service provider route, the distributor's one-month rule, online platforms, how a modification makes you a manufacturer, the heads of damage including data loss, and the exemptions.

10 steps·~15 min
Law & Compliance
advanced

Defectiveness: The Safety a Person Is Entitled to Expect

A product is defective when it lacks the safety a person is entitled to expect. Article 7 turns that into circumstances a court weighs, several written for software: the ability to learn after release, interconnection, cybersecurity requirements, and recalls. This lesson works through the list, the rule that a later improvement is not an admission, and why compliance is not a defence.

10 steps·~15 min
Law & Compliance
advanced

Software as a Product: What the New Liability Directive Changed

Directive (EU) 2024/2853 replaces the 1985 regime and settles a forty-year argument by naming software a product. This lesson covers the new definition and why delivery method is irrelevant, why information is not a product, how components and related services extend the net, where open source sits, and why liability cannot be disclaimed by contract.

10 steps·~15 min