AnyLearn
All lessons
Businessintermediate

Doing It Properly: Vendors, Measurement, and Candidates

The obligations become concrete in three places: what you ask a vendor before buying, what you measure on your own applicants, and what you owe the person a system decided about. This lesson covers all three, plus the AI literacy duty as it applies to recruiters, and the questions worth asking before adopting anything.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

Questions for a vendor

HR tools are almost always bought rather than built, so the diligence conversation is where most of the control sits. Eight questions, and how a vendor answers matters as much as what they say.

Are you the provider of this system under the AI Act, and how have you classified it? A vendor who cannot answer confidently has not done the analysis.

If you say it is not high-risk, which condition of Article 6(3) do you rely on, and how do you address the profiling limit? This is the question that separates vendors who have read the provision from those citing it.

What is the model trained to predict, and what generated that label? If the answer is past hiring decisions or performance ratings, the label problem applies and they should say so.

What is your disaggregated performance, by the groups relevant to our jurisdiction? For a high-risk system, Article 13 requires performance regarding specific groups in the instructions for use, so this is a disclosure rather than a favour.

Can we evaluate on our own applicant data before committing?

What happens to our candidate data, including whether it trains models serving other customers?

How do you notify material model changes?

And what does the candidate see, and what can they challenge?

A vendor claiming their tool is bias-free deserves particular scepticism. It is not a property any system has, and the claim indicates either no measurement or no candour.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. Questions for a vendor

HR tools are almost always bought rather than built, so the diligence conversation is where most of the control sits. Eight questions, and how a vendor answers matters as much as what they say.

Are you the provider of this system under the AI Act, and how have you classified it? A vendor who cannot answer confidently has not done the analysis.

If you say it is not high-risk, which condition of Article 6(3) do you rely on, and how do you address the profiling limit? This is the question that separates vendors who have read the provision from those citing it.

What is the model trained to predict, and what generated that label? If the answer is past hiring decisions or performance ratings, the label problem applies and they should say so.

What is your disaggregated performance, by the groups relevant to our jurisdiction? For a high-risk system, Article 13 requires performance regarding specific groups in the instructions for use, so this is a disclosure rather than a favour.

Can we evaluate on our own applicant data before committing?

What happens to our candidate data, including whether it trains models serving other customers?

How do you notify material model changes?

And what does the candidate see, and what can they challenge?

A vendor claiming their tool is bias-free deserves particular scepticism. It is not a property any system has, and the claim indicates either no measurement or no candour.

2. A vendor's numbers are not your numbers

This point is worth its own step because it is the most common false comfort in HR AI procurement.

A vendor's fairness evaluation was performed on their data: their customers' applicant pools, their labels, their role types, their jurisdictions. Your applicant pool is different in composition, your roles are different, and your historical decisions, which may be what the model was tuned against, are different.

So a vendor's disparity figure tells you the tool can perform acceptably somewhere. It does not tell you how it performs on your applicants, and the deployer's obligation under the Act concerns the system as you use it.

What this means practically. Insist on evaluating with your own data before committing, which the procurement cursus establishes as standard and which HR vendors resist more than most. Hold back a portion the vendor never sees, so you are measuring the tool rather than their tuning against your sample. And measure again after deployment on live outcomes, since the pool that applies to you shifts.

The deployer obligation makes this concrete. Where you control the input data, ensuring it is relevant and sufficiently representative in view of the intended purpose is your duty, not the provider's. Feeding a system an applicant population unlike the one it was validated on is your responsibility.

And in most jurisdictions the employer is the one a claimant sues, whatever the contract says between you and the vendor.

3. What to measure

The measurement an HR function should run, which is a specific instance of the bias cursus applied to a hiring funnel.

Measure at each stage, not only at the end. Application, screening pass, interview invitation, offer, acceptance. A funnel with acceptable overall outcomes can contain a severe drop at one stage that the aggregate hides, and knowing which stage is where the remedy lives.

The metrics. Selection rate by group at each stage, which is the demographic parity view and what the four-fifths screen tests. And where outcome data eventually arrives, the true positive rate by group, which is equal opportunity and speaks to qualified candidates being missed.

Choose the criterion deliberately. In hiring, the false negative, a qualified person excluded from consideration entirely, is usually the harm that matters, which points to equal opportunity.

And compare against the human baseline. The relevant question is whether the tool is better than the process it replaced, measured, not whether it is perfect.

Report the residual on the criteria you did not optimise, since the impossibility result means something always moves.

flowchart LR
A["Applications received"] --> B["Screening pass"]
B --> C["Interview invitation"]
C --> D["Offer"]
D --> E["Acceptance"]
A --> F["Selection rate by group at EACH stage"]
B --> F
C --> F
D --> F
F --> G["Aggregate can hide a severe drop at one stage"]
G --> H["Equal opportunity: qualified candidates missed"]
H --> I["Compare against the human baseline, measured"]

4. The data you need, and may not have

Measuring by group requires knowing the groups, and most HR functions do not hold that data for applicants.

The reasons are understandable. Collecting demographic information from candidates feels intrusive, may be restricted, and creates a data protection obligation. Many organisations deliberately avoid it.

The position has changed, and it is worth knowing. The AI Act permits processing special category data where strictly necessary for bias detection and correction in high-risk systems, subject to safeguards including limits on re-use, pseudonymisation and restricted access, and the 2026 Digital Omnibus clarified and widened that flexibility. So not measuring is no longer defensible on data protection grounds.

Practical routes where you lack the data.

Voluntary self-identification at application, collected separately from the application itself so it cannot influence the decision, with a clear statement of purpose. Response rates are partial and usually sufficient to detect a material disparity.

Measure on the stages where you do have data. Many organisations hold demographic data on hires and not on applicants, which permits some analysis and not the screening-stage question that matters most.

And separate the measurement pipeline from the decision pipeline architecturally, so the attribute is available to the analyst and not to the model.

What is not acceptable is concluding that measurement is impossible and proceeding. The absence of a measurement is not the absence of a disparity, and in a jurisdiction where the burden shifts on a showing of disparate outcomes, it is the absence of a defence.

5. What the candidate is owed

The person the system decided about is the one with the least visibility and the most at stake, and the obligations here are converging across jurisdictions.

That a system was involved. Article 50 transparency duties apply from 2 August 2026 where a system interacts with a person, and beyond the Act, several jurisdictions have introduced or are introducing notice requirements specific to automated employment decisions. Notice is also the precondition for everything else: a decision presented as institutional cannot be challenged as automated.

What it considered. Not the model internals, which nobody can usefully convey, but what the system assessed: experience, qualifications, responses to particular questions.

A route to challenge, with a human who reconsiders and has authority to change the outcome. A challenge returning the same automated answer is not redress, it is the same decision repeated.

And, where applicable, the data protection rights that attach to automated decision-making, including the right to obtain human intervention and to contest the decision.

One practical point that improves outcomes disproportionately. Where a system rejects, saying something specific is far better than a generic notice, both for the candidate and for you: a candidate told the role required a certification they did not list may correct the record, which surfaces exactly the false negatives your measurement is trying to detect.

The redress channel is also a detection route, and in HR it is the only one the affected person controls.

6. Human oversight that is real

Every HR AI deployment claims a human in the loop, and the claim is usually nominal. The four conditions from the regulated-industries cursus apply directly, and Article 26 makes competence, training and authority an obligation of result for high-risk deployers.

Competence. The recruiter can tell whether the system's assessment is wrong, which requires understanding what it measures and where it fails. Generic AI literacy does not establish this for a specific tool.

Authority. Overriding the system is permitted and expected. If advancing a candidate the model ranked low requires written justification while accepting the ranking requires nothing, you have built an incentive gradient and should expect the behaviour it rewards.

Time. The arithmetic is the test. A recruiter reviewing four hundred model-ranked applications in a day has seconds each, which is not review. And the volume argument that justified automation is usually the same pressure that makes oversight impossible, which is a tension worth naming rather than resolving on paper.

Information. The recruiter sees what drove the assessment, not a bare score.

And the diagnostic: track the override rate. A recruiter who advances candidates the model ranked low, sometimes, is exercising judgement. One whose decisions match the ranking almost always is providing a signature, and the model is effectively deciding.

Sample and re-review to distinguish the two, because the number alone cannot.

7. AI literacy for the HR function

Article 4, as amended by the 2026 Digital Omnibus, requires providers and deployers to take appropriate measures to support the development of AI literacy among staff and others operating systems on their behalf. It is now an obligation of effort rather than result, and it does not require guaranteeing any individual's level.

For HR the obligation has a double aspect that is worth noticing.

HR is a deployer, so recruiters and HR staff operating these systems are in scope, and the training should be specific to the tools they use rather than general.

And HR frequently owns the organisation's training function, so it is also the team asked to deliver AI literacy for everyone else. The AI literacy cursus covers designing that programme.

What recruiters specifically need. What the tool assesses and what it does not. Where it performs worse, from the disaggregated measurement. That fluency in output carries no signal about correctness. That their override authority is real and expected. And what to do when a candidate challenges a decision.

What the wider function needs is lighter: what the prohibitions are, so nobody procures an emotion-inference tool believing it to be an engagement analytics product.

And note the neighbouring duty that did not soften. Article 26's requirement that oversight be assigned to persons with the necessary competence is a requirement of result, so for a high-risk deployment, supporting literacy is not sufficient and demonstrable competence is the standard.

8. Five questions before adopting anything

The cursus reduces to a short interrogation worth running on any proposed HR AI deployment.

Does it infer emotion or categorise on biometrics? If yes, stop. This is prohibited and has been since February 2025.

Does it score, rank, flag or filter a person? If yes, it is high-risk in an Annex III area, the derogation almost certainly does not apply because of the profiling limit, and you need the full apparatus: vendor diligence, disaggregated measurement on your own data, designed oversight, candidate notice and a redress route.

What is it trained to predict, and what generated that label? If the answer is past hiring decisions or performance ratings, understand that accuracy means fidelity to past decisions, and say so internally rather than accepting a bias-reduction claim.

What is our current disparity, measured? Without this you cannot tell whether a tool helps, and in a jurisdiction where the burden shifts on disparate outcomes you have no defence either.

And what does the candidate see and what can they do about it? If the answer is nothing and nothing, that is a design decision someone should make consciously.

The closing observation. Most of the realisable value in HR AI sits in drafting, onboarding and administrative support, none of which triggers any of this. The screening tools that attract the most attention carry nearly all of the risk, and the honest reckoning is that they are frequently solving a sourcing problem at the wrong stage.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why is a vendor's fairness evaluation insufficient for your deployment?
    • Vendors are not permitted to share evaluation data
    • It was performed on their applicant pools, labels and jurisdictions, not yours, and the deployer obligation concerns the system as you use it
    • Fairness metrics are not comparable between organisations
    • Vendor evaluations only cover demographic parity
  2. Why should a hiring funnel be measured at each stage rather than end to end?
    • Regulators require stage-level reporting
    • Aggregate figures are statistically unreliable
    • Acceptable overall outcomes can hide a severe drop at one stage, and knowing which stage is where the remedy lives
    • Each stage uses a different model
  3. An organisation says it cannot measure applicant disparities because it lacks demographic data. What is the position?
    • This remains a valid justification under data protection law
    • Only public sector employers may collect such data
    • Proxy inference is the only permitted route
    • The AI Act permits processing special category data for bias detection, so voluntary collection separated from the decision pipeline is available
  4. What indicates that human oversight of a screening tool is nominal?
    • The recruiter's decisions match the model's ranking almost always
    • The recruiter occasionally advances low-ranked candidates
    • The recruiter requests additional information about the model
    • The override rate fluctuates between reviewers
  5. Why is a specific rejection notice better than a generic one, for the employer as well as the candidate?
    • It reduces the volume of challenges received
    • A candidate may correct the record, which surfaces exactly the false negatives your measurement is trying to detect
    • It satisfies the Article 6(3) derogation requirements
    • It transfers liability to the candidate

Related lessons

Law & Compliance
advanced

Proof: Disclosure, Presumptions, and the Complexity Rule

Strict liability is worthless if the claimant cannot prove a defect they never saw. Articles 9 and 10 answer that with a disclosure order, three presumptions of defectiveness, a presumption of causation, and a rule turning complexity into the claimant's ally. This lesson works through the cascade, the three-year and ten-year clocks, and what a defendant should be able to produce.

10 steps·~15 min
Law & Compliance
advanced

Who Pays, and For What Damage

The Directive builds a chain of liable operators so an injured person in the EU always has someone to sue. This lesson covers the manufacturer and component manufacturer, the importer and fulfilment service provider route, the distributor's one-month rule, online platforms, how a modification makes you a manufacturer, the heads of damage including data loss, and the exemptions.

10 steps·~15 min
Law & Compliance
advanced

Defectiveness: The Safety a Person Is Entitled to Expect

A product is defective when it lacks the safety a person is entitled to expect. Article 7 turns that into circumstances a court weighs, several written for software: the ability to learn after release, interconnection, cybersecurity requirements, and recalls. This lesson works through the list, the rule that a later improvement is not an admission, and why compliance is not a defence.

10 steps·~15 min
Law & Compliance
advanced

Software as a Product: What the New Liability Directive Changed

Directive (EU) 2024/2853 replaces the 1985 regime and settles a forty-year argument by naming software a product. This lesson covers the new definition and why delivery method is irrelevant, why information is not a product, how components and related services extend the net, where open source sits, and why liability cannot be disclaimed by contract.

10 steps·~15 min