AIG-025 AI Fairness and Bias Controls
Description
AI systems that score, rank, recommend or classify individuals have documented fairness objectives naming the fairness metric applied and the threshold that counts as a pass, chosen for the system's purpose and the population it affects. A bias evaluation disaggregated by the relevant protected characteristics is recorded before deployment and repeated in production at defined intervals. No system whose evaluation breached its threshold remains deployed without a recorded remediation decision and a re-evaluation result. Evaluation methodology and results are retained and are traceable to the model version they were run against.
Rationale
Bias is not visible in system logs or security tests. It only appears when the outputs are disaggregated and compared against a fairness criterion that was chosen in advance, because a criterion chosen after seeing the numbers can always be met. Production re-evaluation catches the bias that accumulates from feedback loops and data drift rather than from the original training set. AIG-008 holds the verification and validation gate and AIG-026 the adversarial and security evaluation. AIG-012 holds the quality of the dataset the evaluation is run on.
Applicability (9 profiles)
Pre-deployment bias evaluation is the provider's, obtained through AIG-034. Evaluation in production on the deployer's own population is the deployer's.
Art.10(2), points (f) and (g), are the source: the datasets are examined in view of possible biases likely to affect health, safety or fundamental rights or to lead to prohibited discrimination. Appropriate measures to detect, prevent and mitigate those biases are taken. That places the fairness objective, its metric and its threshold inside the data governance an authority reads under Art.11, so they are evidence against Art.10 rather than an internal quality measure. It names the axis the evaluation is disaggregated on. The lawful route to the special categories such an evaluation needs is new Art.4a(1), whose conditions are on AIG-014.
Pre-deployment bias evaluation is the provider's, obtained through AIG-034. Evaluation in production on the deployer's own population is the deployer's.
Framework Mappings (18)
| GRC-11 | Bias and Fairness Assessment | full |
| EU-AI-Art.10.2 | Data Governance — Data Preparation and Bias Management | full |
| EU-AI-Art.15.1 | Accuracy, Robustness and Cybersecurity — Performance Standards | informative |
| GDPR-Art.5.1a | Lawfulness, Fairness and Transparency of Processing | partial |
| MG-2.2-004 | Deployed AI System Value Maintenance | MG-2.2-004 | partial |
| MP-1.2-002 | Interdisciplinary AI Team Composition | MP-1.2-002 | informative |
| MS-1.1-003 | AI Risk Measurement Approach Selection | MS-1.1-003 | partial |
| MS-1.1-006 | AI Risk Measurement Approach Selection | MS-1.1-006 | full |
| MS-1.3-001 | Independent AI Risk Assessment | MS-1.3-001 | partial |
| MS-2.11-001 | AI Fairness and Bias Evaluation | MS-2.11-001 | partial |
| MS-2.11-002 | AI Fairness and Bias Evaluation | MS-2.11-002 | full |
| MS-2.11-004 | AI Fairness and Bias Evaluation | MS-2.11-004 | informative |
| MS-2.13-001 | Measurement Effectiveness Evaluation | MS-2.13-001 | informative |
| MS-2.2-001 | Human Subject Evaluation Requirements | MS-2.2-001 | partial |
| MS-3.3-003 | User and Community Feedback Processes | MS-3.3-003 | partial |
| GOVERN 3.1 | Diverse Team Decision-Making | partial |
| MAP 1.2 | Interdisciplinary AI Team Composition | informative |
| MEASURE 2.11 | AI Fairness and Bias Evaluation | full |
Evidence (2)
Bias evaluation report produced before deployment and periodically in production, disaggregated by relevant protected characteristics, with results compared to documented fairness thresholds.
Example: Bias Evaluation Report · Loan Scoring Model v4.1 (Weights & Biases artefact, 2026-Q1), showing demographic parity difference ≤ 0.05 for gender and ethnicity, equalised odds gap ≤ 0.03, comparison to thresholds defined in AI fairness objectives, result: PASS
Test: Request bias evaluation reports for a sample of AI systems subject to bias risk, covering pre-deployment and at least one in-production evaluation. Verify: (1) evaluation is disaggregated by relevant protected characteristics for the system's context, (2) fairness metric definitions match those documented in the AI fairness objectives, (3) results are compared to documented pass/fail thresholds, (4) threshold failures have a documented remediation action and re-evaluation result, (5) production evaluation frequency matches the defined schedule.
Fairness evaluation tool output carrying the disaggregated metrics for each production model, linked to the model version evaluated.
Example: Fairness evaluation run, hiring-screen v2.4, 2026-05-18: demographic parity difference and equalised odds difference by protected characteristic, threshold 0.05, result PASS
Test: Run or retrieve the fairness evaluation output for each model in scope. Verify: (1) the output is disaggregated by the protected characteristics the fairness objectives name, (2) each metric is reported against the threshold the objectives set, (3) the run identifies the model version it evaluated and that version matches the deployed one, (4) a subgroup too small to evaluate is reported as such rather than dropped silently, (5) a run that breached its threshold is followed by a re-evaluation after the recorded remediation.
Questions (2)
Do AI systems that score, rank, recommend or classify individuals have documented fairness objectives?
Bias cannot be detected without predefined, measurable fairness criteria. Objectives should specify the fairness metric (e.g. demographic parity, equalised odds) and the pass/fail threshold, relative to the system's purpose and affected population.
How is bias testing conducted for AI systems subject to bias risk in your organisation?
All five practices are expected. Bias testing conducted only at deployment without production monitoring misses in-production bias accumulation from feedback loops and data drift.