AIG-008 AI System Verification, Validation and Testing
Description
Defined verification and validation (V&V) procedures are executed before an AI system is deployed or after a substantial modification. Testing includes: functional accuracy against defined metrics, performance against pre-specified thresholds, robustness to distributional shift, safety and failure-mode testing, and fairness and bias evaluation across relevant population subgroups. Test datasets, metrics, tooling, and results are documented and retained. Testing is not performed solely by the team that built the system.
Rationale
AI systems fail in qualitatively different ways from conventional software; V&V must be designed specifically for AI failure modes, not inherited from generic software testing.
Applicability (9 profiles)
Obtain the provider's verification and validation summary (AIG-034). The deployer's acceptance checks in its own context are the pre-deployment checks in AIG-009 and, where the system scores people, the in-production evaluation in AIG-025.
Art.9(6) and (8) fix what the gate measures against: testing throughout development and in any event before the system is placed on the market or put into service, against pre-defined metrics and probabilistic thresholds, to identify the risk management measures and to verify compliance with Section 2. The threshold is therefore set before the test rather than read off it. Art.60 adds a second regime where the testing runs in real-world conditions outside a regulatory sandbox and it is a permission rather than a practice: a real-world testing plan submitted to the market surveillance authority of the Member State the testing runs in, that authority's approval, with tacit approval after 30 days only where national law provides for it, registration under a Union-wide single identification number with the Annex IX information and a provider established in the Union or holding a legal representative who is.
Obtain the provider's verification and validation summary (AIG-034). The deployer's acceptance checks in its own context are the pre-deployment checks in AIG-009 and, where the system scores people, the in-production evaluation in AIG-025.
Framework Mappings (38)
| MDS-07 | Robustness against Adversarial Attack / Model Hardening | informative |
| MDS-11 | Model Failure | partial |
| EU-AI-Art.15.1 | Accuracy, Robustness and Cybersecurity — Performance Standards | partial |
| EU-AI-Art.60 | Real-World Testing — Plan, Registration and Informed Consent | partial |
| EU-AI-Art.9.5 | AI Risk Management System — Testing for Risk Management | full |
| COP-S-7.3 | Documentation of systemic risk identification, analysis, and mitigation | informative |
| A.6.2.4 | AI system verification and validation | full |
| GV-1.2-002 | Trustworthy AI Characteristics Integration | GV-1.2-002 | partial |
| GV-1.3-002 | Risk Management Activity Level Determination | GV-1.3-002 | full |
| GV-1.5-003 | Risk Management Monitoring and Review | GV-1.5-003 | informative |
| GV-3.2-001 | Human-AI Configuration Roles | GV-3.2-001 | partial |
| GV-4.1-002 | Safety-First Organisational Culture | GV-4.1-002 | partial |
| MP-2.1-002 | AI System Task and Method Definition | MP-2.1-002 | partial |
| MP-2.3-001 | Scientific Integrity and Testing Considerations | MP-2.3-001 | informative |
| MP-3.4-004 | Operator Proficiency Processes | MP-3.4-004 | informative |
| MP-3.4-006 | Operator Proficiency Processes | MP-3.4-006 | partial |
| MP-4.1-007 | AI Technology and Legal Risk Mapping | MP-4.1-007 | informative |
| MP-5.1-005 | Impact Likelihood and Magnitude Documentation | MP-5.1-005 | informative |
| MS-1.3-002 | Independent AI Risk Assessment | MS-1.3-002 | partial |
| MS-1.3-003 | Independent AI Risk Assessment | MS-1.3-003 | partial |
| MS-2.11-001 | AI Fairness and Bias Evaluation | MS-2.11-001 | informative |
| MS-2.11-002 | AI Fairness and Bias Evaluation | MS-2.11-002 | informative |
| MS-2.13-001 | Measurement Effectiveness Evaluation | MS-2.13-001 | partial |
| MS-2.3-001 | AI System Performance Measurement | MS-2.3-001 | informative |
| MS-2.3-002 | AI System Performance Measurement | MS-2.3-002 | full |
| MS-2.3-004 | AI System Performance Measurement | MS-2.3-004 | partial |
| MS-2.5-001 | AI System Validity and Reliability | MS-2.5-001 | full |
| MS-3.3-003 | User and Community Feedback Processes | MS-3.3-003 | informative |
| GOVERN 4.3 | AI Testing and Information Sharing Practices | informative |
| MEASURE 1.1 | AI Risk Measurement Approach Selection | partial |
| MEASURE 1.3 | Independent AI Risk Assessment | full |
| MEASURE 2.1 | AI Testing and Evaluation Documentation | full |
| MEASURE 2.11 | AI Fairness and Bias Evaluation | full |
| MEASURE 2.3 | AI System Performance Measurement | full |
| MEASURE 2.5 | AI System Validity and Reliability | full |
| MEASURE 2.6 | AI System Safety Risk Evaluation | full |
| MEASURE 4.1 | Deployment-Context Risk Measurement | partial |
| MEASURE 4.2 | Trustworthiness Measurement with Expert Input | partial |
Evidence (2)
AI system V&V test report produced before deployment or after substantial modification, documenting test datasets, metrics, tooling, results, and independent reviewer sign-off.
Example: V&V Test Report · Fraud Detection Model v4 (Confluence), dated 2025-11-03, containing accuracy, precision, recall, fairness metrics, adversarial robustness results, and sign-off by independent QA team
Test: Request the V&V test report for a sample of production AI systems. Verify: (1) report covers functional accuracy, robustness, safety, and fairness dimensions, (2) test datasets are identified and versioned, (3) results are compared to pre-specified acceptance thresholds, (4) report was authored or reviewed by a team independent of the development team, (5) report is retained and accessible.
Automated evaluation pipeline output (CI/CD test suite results) demonstrating that defined model performance thresholds were checked programmatically before promotion to production.
Example: GitHub Actions CI pipeline run log for model-fraud-detection (run #4812), showing automated accuracy >= 0.92, AUC >= 0.95, and bias test pass gates before merge approval
Test: Request CI/CD pipeline logs for a recent model deployment. Verify: (1) automated evaluation steps are present in the pipeline definition, (2) performance thresholds are coded as pass/fail gates, (3) the deployment was blocked or approved based on gate results, (4) bias evaluation is included as a gate (not only accuracy).
Questions (2)
Are defined verification and validation procedures executed before any AI system is deployed or after a substantial modification?
AI V&V must cover dimensions conventional software testing misses: distributional robustness, fairness across population subgroups, and safety failure modes. Testing should not be performed solely by the team that built the system.
Which of the following are included in your AI system V&V testing?
All six elements characterise a mature AI V&V process. Fairness evaluation and independent review are the most frequently absent from programmes that inherit generic software testing practices.