AnyLearn
All lessons
Businessintermediate

What You Are Actually Buying: Scoping an AI Purchase

AI procurement fails at the scoping stage, before any vendor is contacted. This lesson covers what makes an AI purchase different from ordinary software, the regulatory position you inherit from the seller, the questions that determine whether you become a provider yourself, how to specify a problem rather than a product, and the build-buy-or-do-nothing decision that should precede any shortlist.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 10

Why AI procurement is different

Buying ordinary software is a question about features, price and support. Four things make an AI purchase structurally different, and each of them is invisible in a demo.

The product's behaviour is statistical. It has an error rate, that rate differs across the population it is used on, and neither number appears on a pricing page. Two tools with identical feature lists can differ enormously in the thing that matters.

The performance you observe is conditional on data you have not seen. A vendor's accuracy claim was established on some distribution. Whether it holds on yours is an empirical question, not a contractual one.

The purchase changes your regulatory position. Buying a tool can make you a deployer with real obligations, and configuring it a certain way can make you a provider with far heavier ones.

And the thing you buy does not stay the same. The vendor updates the model, behaviour shifts, and nothing was deployed on your side.

Ordinary procurement processes catch none of these, which is why AI purchases that passed a normal review still produce unpleasant surprises.

Full lesson text

All 10 steps on one page, for reading, reference, and search.

Show

1. Why AI procurement is different

Buying ordinary software is a question about features, price and support. Four things make an AI purchase structurally different, and each of them is invisible in a demo.

The product's behaviour is statistical. It has an error rate, that rate differs across the population it is used on, and neither number appears on a pricing page. Two tools with identical feature lists can differ enormously in the thing that matters.

The performance you observe is conditional on data you have not seen. A vendor's accuracy claim was established on some distribution. Whether it holds on yours is an empirical question, not a contractual one.

The purchase changes your regulatory position. Buying a tool can make you a deployer with real obligations, and configuring it a certain way can make you a provider with far heavier ones.

And the thing you buy does not stay the same. The vendor updates the model, behaviour shifts, and nothing was deployed on your side.

Ordinary procurement processes catch none of these, which is why AI purchases that passed a normal review still produce unpleasant surprises.

2. Start with the decision, not the tool

The most expensive procurement mistakes are made before a vendor is contacted, in how the requirement is framed.

A requirement stated as we need an AI tool for recruitment has already made several decisions invisibly: that the problem is a tooling problem, that AI is the right instrument, and that the scope is the whole of recruitment.

Restate it as a decision. Which specific decision is this system going to inform or make? Who is affected by that decision? What does the current process get wrong, and how would we know if the new one got it wrong less?

That reframing does three useful things.

It reveals the regulatory position immediately, because the risk tier follows the decision, not the product. A tool informing who gets interviewed is in Annex III territory before you have spoken to anyone.

It produces an evaluation criterion. If you cannot say what the current process gets wrong, you cannot tell whether a replacement is better, and every vendor will look convincing.

And it surfaces the option that vendors will not raise: fixing the process without buying anything. A screening bottleneck caused by an unclear job description is not solved by ranking applicants faster.

3. The role question, answered before shortlisting

Whether you end up a deployer or a provider is decided by how you intend to use the system, and it should be settled before evaluation rather than discovered after deployment.

You remain a deployer if you use the system for the purpose the vendor declared, under the vendor's name, without modifying it.

You may become the provider, taking on the full set of provider obligations, in three situations covered by the Act: you put your own name or trademark on a high-risk system already on the market; you substantially modify such a system; or you modify the intended purpose of a system, including a general-purpose AI system, so that it becomes high-risk.

Three procurement patterns walk into that third route without noticing.

White-labelling a vendor's tool so it appears to customers as your product.

Buying a general-purpose platform and building a workflow on it that makes a decision about people. The platform vendor did not put a recruitment system on the market. You did.

Heavy configuration that moves the system beyond the declared intended purpose, which is a spectrum rather than a line and therefore needs a documented judgement.

Settle this at scoping. The cost difference between the two roles is larger than any difference between vendors on the shortlist.

4. The scoping gate

A short sequence that runs before any vendor conversation, and which determines how much diligence the purchase warrants.

Start from the decision the system will inform, not from the product category. Establish who is affected: staff, customers, the public, or nobody outside the organisation.

Where individuals are affected, check the Annex III areas. That determines whether this is an ordinary software purchase or one carrying substantial obligations.

Then the role question: will you use it as supplied, or brand, modify or repurpose it?

The output of the gate is a diligence tier. A minimal-risk tool used as supplied deserves a light review. A high-risk system, or any purchase where you become the provider, deserves the full process covered in the next lesson.

The point of the gate is to spend diligence where it changes the answer, rather than applying one process to every purchase and exhausting everyone before the important one arrives.

flowchart TD
A["What decision will this system inform?"] --> B["Who is affected by that decision?"]
B --> C["Nobody outside the organisation: light review"]
B --> D["Individuals affected: check Annex III areas"]
D --> E["Outside Annex III: standard diligence"]
D --> F["Inside Annex III: high-risk diligence"]
A --> G["Will we brand, modify or repurpose it?"]
G --> H["No: we are a deployer"]
G --> I["Yes: we may become the provider"]
I --> J["Full provider obligations: reassess build vs buy"]

5. Build, buy, or neither

The regulatory position changes the classic build-versus-buy calculation in a direction that is not obvious.

Buying a high-risk system from an established provider transfers the heaviest obligations to them. They carry the risk management system, the data governance work, the technical documentation, the conformity assessment, the CE marking, the registration and the post-market monitoring. You carry the deployer set, which is real but far smaller.

Building the same system in-house makes you the provider, and all of that becomes yours. For an organisation without an existing quality management system, this is the single largest undertaking in the Act.

So for high-risk use cases, the regulatory arithmetic favours buying more strongly than the pure engineering arithmetic does, and teams that would happily build a scoring model in a fortnight should price the compliance apparatus before choosing.

The third option deserves equal standing on the paper. Not automating the decision at all, or automating a narrower part of it that stays outside Annex III, is frequently the best answer and almost never appears in a vendor-led process.

A useful discipline: write down what happens if you do nothing, with the same rigour as the vendor options. If the do-nothing case is genuinely unacceptable, the comparison is stronger. If it is not, you have saved a procurement.

6. What the vendor's position tells you

Before any technical evaluation, establish where the vendor sits regulatorily. The answers are diagnostic well beyond compliance.

Are you the provider of this system under the AI Act? A vendor that cannot answer confidently has not done the analysis, which tells you something about the rest of their preparation.

How have you classified it, and on what basis? Ask for the reasoning, not the conclusion. A vendor claiming a recruitment tool is not high-risk should be able to explain which condition of the Article 6(3) derogation they rely on and how they handle the profiling limit.

If you rely on the derogation, is the system registered, and can we see the documented assessment? Registration applies even under the derogation, so a vendor unaware of that has not read the provision they are relying on.

Was the system placed on the EU market before the applicable high-risk date? This determines whether grandfathering applies, and grandfathering means the system may not have been through the requirements at all.

And what is your position on substantial modification, and what change envelope did you declare?

These questions cost nothing and separate vendors who have done the work from those who added a compliance page to their website.

7. Specifying performance before you look at products

Write down what adequate performance means before seeing any vendor's numbers. Doing it afterwards means calibrating your standard to what is available, which is how organisations end up satisfied with tools that do not work well enough.

Four things to specify.

The metric that matters for this decision, and why. For a screening tool that filters, false negatives, meaning qualified people wrongly excluded, are usually the costly error, and they are the error that aggregate accuracy hides and that vendors rarely report.

The population breakdown you require results for. Performance on the groups your decision affects, not just overall. This is the number that will not be volunteered.

The threshold below which you would not proceed, set before you know what is achievable.

The conditions under which the claim must hold: your data, your geography, your volume, your edge cases.

Then ask each vendor for results against that specification rather than accepting their materials. Vendors who cannot produce disaggregated performance either have not measured it or have measured it and would rather not say, and both are informative.

This specification is also what makes the pilot in the next lesson meaningful, since a pilot without a prior threshold is a demonstration rather than a test.

8. Data questions that decide the purchase

Where your data goes is frequently the question that eliminates a vendor, and it is separable from performance.

What happens to input data? Is it retained, for how long, and where. Is it used to train or improve models, for you alone or across customers. Whether a training opt-out exists, whether it is on by default, and whether it applies retrospectively.

Who can access it? Subprocessors, jurisdictions, support staff, and under what controls.

What leaves your boundary? For a system built on a third-party model, your input may reach that model provider, so the answer involves a party you are not contracting with.

And what happens at the end? Deletion on termination, in what timeframe, with what confirmation, including from any model trained on your data, which is frequently impossible and should be established as impossible before rather than after.

Two cautions. A vendor's public documentation and their enterprise contract often differ, so ask for the contractual position. And where the system processes personal data, this analysis has a GDPR half that should be run at the same time by the same people, since discovering an international transfer problem after selecting a vendor restarts the procurement.

9. Reading a demo honestly

Demos are constructed, and knowing how they are constructed makes them useful rather than misleading.

A demo runs on data the vendor chose, on cases where the system performs, with a presenter who knows which inputs work. None of that is dishonest. It is simply not evidence about your data.

Four moves make a demo informative.

Bring your own inputs, including ones you know are hard: the ambiguous case, the unusual format, the edge case that trips your current process. A vendor unwilling to run unseen inputs has told you something.

Ask to see it fail. A vendor who can describe where their system performs poorly is more credible than one who cannot, and the answer tells you whether they have measured.

Watch the interface for the oversight affordances the Act requires of high-risk systems: does it surface confidence, expose what drove the output, and make disagreeing with it easy? A tool that presents a bare score with no way to interrogate it will make meaningful human oversight difficult regardless of what your process says.

And notice what the demo skips. Set-up, integration, the exception path, and what happens when the system is unavailable are usually where the real cost is.

10. The scoping output

By the end of scoping, before a shortlist exists, five things should be written down.

The decision the system will inform, stated specifically enough that a change of use would be visible.

Who is affected, and whether the use case sits in an Annex III area.

Your intended role, deployer or provider, and the reasoning if the answer is uncomfortable.

The performance specification: metric, population breakdown, threshold, and the conditions under which it must hold.

The do-nothing case, assessed honestly.

This document does three jobs at once. It is the brief you give vendors, so responses are comparable. It is the input to your inventory entry and classification, so the governance framework gets fed as a byproduct rather than as a later exercise. And it is the record of reasoning that the AI Act's judgement calls require, written before the answer was known, which is worth considerably more than the same reasoning reconstructed afterwards.

With the scope settled, the next lesson covers the diligence and contract terms that turn it into a purchase you can defend.

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. Why should an AI requirement be framed as a decision rather than a tool?
    • Because vendors will not respond to product-based briefs
    • Because the risk tier follows the decision, and the framing also produces an evaluation criterion and surfaces the do-nothing option
    • Because the AI Act requires decisions to be documented in this form
    • Because it reduces the number of vendors that need to be contacted
  2. Which procurement pattern most commonly turns a buyer into a provider without anyone intending it?
    • Negotiating a discount on an enterprise licence
    • Using a vendor's tool exactly as its documentation describes
    • Building a workflow on a general-purpose platform that makes a decision about people
    • Asking the vendor for disaggregated performance data
  3. How does the AI Act change the build-versus-buy calculation for a high-risk use case?
    • It favours buying, because an established provider carries the conformity assessment, documentation and monitoring obligations
    • It favours building, because in-house systems are exempt from the requirements
    • It makes no difference, since obligations attach to the system rather than the party
    • It requires all high-risk systems to be purchased from certified vendors
  4. Why should the performance threshold be set before looking at vendor materials?
    • Because vendors are contractually required to meet published thresholds
    • Because thresholds set afterwards get calibrated to what happens to be available
    • Because the AI Act specifies minimum accuracy levels
    • Because it shortens the procurement timeline
  5. What does a vendor's inability to produce disaggregated performance results indicate?
    • That the system is exempt from Article 10 data governance requirements
    • That the system is not high-risk
    • That the metric chosen was inappropriate for the use case
    • Either that they have not measured it, or that they have and would rather not say

Related lessons

Law & Compliance
advanced

Proof: Disclosure, Presumptions, and the Complexity Rule

Strict liability is worthless if the claimant cannot prove a defect they never saw. Articles 9 and 10 answer that with a disclosure order, three presumptions of defectiveness, a presumption of causation, and a rule turning complexity into the claimant's ally. This lesson works through the cascade, the three-year and ten-year clocks, and what a defendant should be able to produce.

10 steps·~15 min
Law & Compliance
advanced

Who Pays, and For What Damage

The Directive builds a chain of liable operators so an injured person in the EU always has someone to sue. This lesson covers the manufacturer and component manufacturer, the importer and fulfilment service provider route, the distributor's one-month rule, online platforms, how a modification makes you a manufacturer, the heads of damage including data loss, and the exemptions.

10 steps·~15 min
Law & Compliance
advanced

Defectiveness: The Safety a Person Is Entitled to Expect

A product is defective when it lacks the safety a person is entitled to expect. Article 7 turns that into circumstances a court weighs, several written for software: the ability to learn after release, interconnection, cybersecurity requirements, and recalls. This lesson works through the list, the rule that a later improvement is not an admission, and why compliance is not a defence.

10 steps·~15 min
Law & Compliance
advanced

Software as a Product: What the New Liability Directive Changed

Directive (EU) 2024/2853 replaces the 1985 regime and settles a forty-year argument by naming software a product. This lesson covers the new definition and why delivery method is irrelevant, why information is not a product, how components and related services extend the net, where open source sits, and why liability cannot be disclaimed by contract.

10 steps·~15 min