GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-008 AI System Verification, Validation and Testing

Tier 2+AIProviderDeployerGPAI Model ProviderManaged Service Provider

Description

Defined verification and validation (V&V) procedures are executed before an AI system is deployed or after a substantial modification. Testing includes: functional accuracy against defined metrics, performance against pre-specified thresholds, robustness to distributional shift, safety and failure-mode testing, and fairness and bias evaluation across relevant population subgroups. Test datasets, metrics, tooling, and results are documented and retained. Testing is not performed solely by the team that built the system.

Rationale

AI systems fail in qualitatively different ways from conventional software; V&V must be designed specifically for AI failure modes, not inherited from generic software testing.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredsatisfied by provider

Obtain the provider's verification and validation summary (AIG-034). The deployer's acceptance checks in its own context are the pre-deployment checks in AIG-009 and, where the system scores people, the in-production evaluation in AIG-025.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.9(6) and (8) fix what the gate measures against: testing throughout development and in any event before the system is placed on the market or put into service, against pre-defined metrics and probabilistic thresholds, to identify the risk management measures and to verify compliance with Section 2. The threshold is therefore set before the test rather than read off it. Art.60 adds a second regime where the testing runs in real-world conditions outside a regulatory sandbox and it is a permission rather than a practice: a real-world testing plan submitted to the market surveillance authority of the Member State the testing runs in, that authority's approval, with tacit approval after 30 days only where national law provides for it, registration under a Union-wide single identification number with the Annex IX information and a provider established in the Union or holding a legal representative who is.

Public Body Deployer (EU)stablerequiredsatisfied by provider

Obtain the provider's verification and validation summary (AIG-034). The deployer's acceptance checks in its own context are the pre-deployment checks in AIG-009 and, where the system scores people, the in-production evaluation in AIG-025.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (38)

MDS-07Robustness against Adversarial Attack / Model Hardeninginformative
MDS-11Model Failurepartial
EU-AI-Art.15.1Accuracy, Robustness and Cybersecurity — Performance Standardspartial
EU-AI-Art.60Real-World Testing — Plan, Registration and Informed Consentpartial
EU-AI-Art.9.5AI Risk Management System — Testing for Risk Managementfull
COP-S-7.3Documentation of systemic risk identification, analysis, and mitigationinformative
A.6.2.4AI system verification and validationfull
GV-1.2-002Trustworthy AI Characteristics Integration | GV-1.2-002partial
GV-1.3-002Risk Management Activity Level Determination | GV-1.3-002full
GV-1.5-003Risk Management Monitoring and Review | GV-1.5-003informative
GV-3.2-001Human-AI Configuration Roles | GV-3.2-001partial
GV-4.1-002Safety-First Organisational Culture | GV-4.1-002partial
MP-2.1-002AI System Task and Method Definition | MP-2.1-002partial
MP-2.3-001Scientific Integrity and Testing Considerations | MP-2.3-001informative
MP-3.4-004Operator Proficiency Processes | MP-3.4-004informative
MP-3.4-006Operator Proficiency Processes | MP-3.4-006partial
MP-4.1-007AI Technology and Legal Risk Mapping | MP-4.1-007informative
MP-5.1-005Impact Likelihood and Magnitude Documentation | MP-5.1-005informative
MS-1.3-002Independent AI Risk Assessment | MS-1.3-002partial
MS-1.3-003Independent AI Risk Assessment | MS-1.3-003partial
MS-2.11-001AI Fairness and Bias Evaluation | MS-2.11-001informative
MS-2.11-002AI Fairness and Bias Evaluation | MS-2.11-002informative
MS-2.13-001Measurement Effectiveness Evaluation | MS-2.13-001partial
MS-2.3-001AI System Performance Measurement | MS-2.3-001informative
MS-2.3-002AI System Performance Measurement | MS-2.3-002full
MS-2.3-004AI System Performance Measurement | MS-2.3-004partial
MS-2.5-001AI System Validity and Reliability | MS-2.5-001full
MS-3.3-003User and Community Feedback Processes | MS-3.3-003informative
GOVERN 4.3AI Testing and Information Sharing Practicesinformative
MEASURE 1.1AI Risk Measurement Approach Selectionpartial
MEASURE 1.3Independent AI Risk Assessmentfull
MEASURE 2.1AI Testing and Evaluation Documentationfull
MEASURE 2.11AI Fairness and Bias Evaluationfull
MEASURE 2.3AI System Performance Measurementfull
MEASURE 2.5AI System Validity and Reliabilityfull
MEASURE 2.6AI System Safety Risk Evaluationfull
MEASURE 4.1Deployment-Context Risk Measurementpartial
MEASURE 4.2Trustworthiness Measurement with Expert Inputpartial

Evidence (2)

reportdocumentmanual

AI system V&V test report produced before deployment or after substantial modification, documenting test datasets, metrics, tooling, results, and independent reviewer sign-off.

Example: V&V Test Report · Fraud Detection Model v4 (Confluence), dated 2025-11-03, containing accuracy, precision, recall, fairness metrics, adversarial robustness results, and sign-off by independent QA team

Test: Request the V&V test report for a sample of production AI systems. Verify: (1) report covers functional accuracy, robustness, safety, and fairness dimensions, (2) test datasets are identified and versioned, (3) results are compared to pre-specified acceptance thresholds, (4) report was authored or reviewed by a team independent of the development team, (5) report is retained and accessible.

tool_outputtechnicalautomated

Automated evaluation pipeline output (CI/CD test suite results) demonstrating that defined model performance thresholds were checked programmatically before promotion to production.

Example: GitHub Actions CI pipeline run log for model-fraud-detection (run #4812), showing automated accuracy >= 0.92, AUC >= 0.95, and bias test pass gates before merge approval

Test: Request CI/CD pipeline logs for a recent model deployment. Verify: (1) automated evaluation steps are present in the pipeline definition, (2) performance thresholds are coded as pass/fail gates, (3) the deployment was blocked or approved based on gate results, (4) bias evaluation is included as a gate (not only accuracy).

Questions (2)

boolean

Are defined verification and validation procedures executed before any AI system is deployed or after a substantial modification?

AI V&V must cover dimensions conventional software testing misses: distributional robustness, fairness across population subgroups, and safety failure modes. Testing should not be performed solely by the team that built the system.

multi

Which of the following are included in your AI system V&V testing?

Functional accuracy against pre-specified metrics and thresholdsRobustness to distributional shift or out-of-distribution inputsSafety and failure-mode testingFairness and bias evaluation across protected characteristic subgroupsIndependent review (not solely by the development team)Documented and retained test datasets and resultsNone of the above

All six elements characterise a mature AI V&V process. Fairness evaluation and independent review are the most frequently absent from programmes that inherit generic software testing practices.