GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-022 Human Oversight of AI Outputs

Tier 2+AIProviderDeployerGPAI Model ProviderManaged Service Provider

Description

AI systems whose outputs are used in decisions affecting individuals have a documented oversight design naming, for each system, the oversight measure in force, whether it applies before the output is acted on or after, the point in the process at which it applies and the role that performs it. The design states the competencies the oversight role requires, the training it receives including automation bias and the time allowed for each review. Oversight activity is recorded with the reviewer, the timestamp and the decision. The override rate is monitored against the range the design states as expected.

Rationale

Oversight is the last barrier between a wrong output and a consequence for a person. It also fails quietly: a reviewer with sixty cases an hour and a default of agree produces a complete audit trail and no oversight at all. That is why the time allowance and the override rate are in the control. An override rate at or near zero over a long period is the signal to investigate. Whether review must precede the action is an assurance-depth decision carried by the tier model and the profile, not by this text. AIG-023 holds the override and deactivation mechanisms this oversight relies on and AIG-017 the explanation the reviewer reads.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredrole duty

Assignment of oversight to competent, trained persons and the record of oversight activity are the deployer's (Art.26.2). The oversight measures available are designed by the provider (Art.14) and described in the instructions for use.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.14(1) and (3) move oversight upstream from an operating procedure to a design obligation: the system is designed and developed, human-machine interface tools included, so that natural persons can effectively oversee it while it is in use, with the measures commensurate with the risks, the level of autonomy and the context. The oversight design is therefore drawn before the system is placed on the market and travels with it and Art.14(4) fixes the capabilities it has to deliver to the person holding oversight. For remote biometric identification under Annex III, point 1(a), Art.14(5) requires that no action or decision is taken on an identification result unless at least two natural persons with the necessary competence, training and authority have separately verified and confirmed it (extract row EU-AI-Art.14.3); the deployer performs that verification and the provider designs the system so it can be performed.

Public Body Deployer (EU)stablerequiredrole duty

Art.26(2) binds both seats: oversight is assigned to natural persons who have the competence, the training and the authority to perform it, with the support resources to do so. Art.14(5) adds a step the provider designs and the deployer performs: no action or decision is taken on an identification result produced by a remote biometric identification system under Annex III point 1(a) unless at least two natural persons with the necessary competence, training and authority have separately verified and confirmed it. The extract row, EU-AI-Art.14.3, maps partial to this control since migration 063 with the two-person step named as the gap, so the oversight design states that step for any such system. HRS-013 carries the competence the persons need.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (18)

GRC-15Human supervisionpartial
IAM-17Output Modification and Special Authorizationpartial
EU-AI-Art.14.1Human Oversight — System Design for Oversightfull
EU-AI-Art.14.2Human Oversight — Capabilities Assigned to Oversight Personsfull
EU-AI-Art.14.3Human Oversight — Dual Verification for Biometric Identificationpartial
EU-AI-Art.26.2Deployer Obligations — Human Oversight Assignmentfull
EU-AI-Art.5.3Prohibited Practices — Prior Authorisation of Each Real-Time Remote Biometric Identification Useinformative
GDPR-Art.22Automated Decision-Making and Profilinginformative
AML.M0029Human In-the-Loop for AI Agent Actionsinformative
MG-3.2-008Pre-Trained Model Monitoring | MG-3.2-008full
MP-3.4-005Operator Proficiency Processes | MP-3.4-005full
MS-3.3-002User and Community Feedback Processes | MS-3.3-002informative
MS-4.2-004Trustworthiness Measurement with Expert Input | MS-4.2-004full
GOVERN 3.2Human-AI Configuration Rolesfull
MAP 3.4Operator Proficiency Processesinformative
MAP 3.5Human Oversight Process Definitionfull
ASI09Human-Agent Trust Exploitationpartial
LLM07Misinformationinformative

Evidence (2)

recorddocumentmanual

Oversight design or operational procedure for each production AI system whose outputs are used in decisions affecting individuals, naming the oversight measure in force, the role that performs it, the competencies required, the training given and the time allowed per review.

Example: Human Oversight Procedure · AI Credit Decisioning System (Confluence), specifying that all AI-flagged decline decisions require human review within 4 hours, reviewer qualification requirements (credit underwriting certification), automation bias awareness training requirement, and escalation path for reviewer disagreement

Test: Request the oversight design for each production AI system whose outputs are used in decisions affecting individuals. Verify: (1) the oversight measure in force is described for that system, (2) the design states whether review precedes the action or follows it and names the point in the process at which it applies, (3) competency requirements for the oversight role are defined, (4) automation bias is addressed in the training or the operator guidance, (5) the time allowed per review is stated and is consistent with the volume of cases the role receives.

logtechnicalautomated

Audit trail records showing human review and override events for AI-generated outputs, demonstrating that oversight is operationally active and not merely nominal.

Example: AI-Credit-System override log (Splunk, last 90 days): 1,247 AI decisions reviewed, 89 overrides recorded with reviewer ID, timestamp, and override reason category; override rate 7.1%, consistent with expected range 5–10%

Test: Request the oversight design and the override and review event logs for a 90-day sample. Verify: (1) every review event records the reviewer, the timestamp and the decision, (2) every override event records a reason category, (3) the oversight design states the override range it expects for the system, (4) the measured override rate over the sample is compared against that range and a rate outside it carries a recorded investigation, (5) the review volume in the log matches the population of outputs the design says are reviewed.

Questions (2)

boolean

Do AI systems that produce outputs used in decisions affecting individuals have documented human oversight mechanisms proportionate to their risk level?

Human oversight is the last line of defence against harmful AI outputs. It must be substantively designed, not nominal. Oversight persons must have defined competencies, training, and sufficient time to conduct meaningful review.

multi

Which of the following are true of the human oversight of your AI systems used in decisions affecting individuals?

The oversight design is documented and specific to each systemThe design states whether review takes place before the output is acted on or afterThe oversight role has defined competency requirementsTraining for the oversight role covers automation biasThe design states the time allowed for each reviewOverride and escalation paths are documented and accessible to the reviewerOverride rates are monitored against an expected rangeNone of the above

Every item applies to each system in scope. An override rate at or near zero sustained over a long period is a signal that review is nominal rather than substantive, which is why monitoring the rate against an expected range matters more than the rate itself.