The question behind every question
A supervisor examining an AI system is not primarily asking whether it works. They are asking how you know, and whether anyone independent of the people who built it checked.
That reframing explains most of what follows. Controls that would satisfy an engineering team, we tested it and the numbers were good, do not satisfy a supervisor, because the assertion and the evidence come from the same source. What is being assessed is not the system alone but the organisation's ability to have caught a problem.
The corollary is uncomfortable and useful: a well-performing system with no evidence trail is in a worse regulatory position than a mediocre one that was properly validated, monitored and documented. Regulators cannot verify performance directly. They verify process, and infer.
This is not bureaucratic obtuseness. It is the only workable approach when the supervised population is large, the systems are opaque, and the supervisor arrives after the fact.
The rest of this cursus is about producing that evidence for systems whose behaviour is statistical, whose inputs are unbounded, and which change after deployment. All three properties make the established machinery fit awkwardly, and the awkwardness is the subject.

