AIG-013 Training Data Provenance
Description
The provenance of every dataset used to train, fine-tune or evaluate a production model is recorded. The record names the origin of each source, whether internal, third-party, licensed, web-sourced or synthetic, its licence and copyright status, the date it was collected or acquired, the transformations applied to it and the lawful basis relied on for any personal data it holds. Each provenance record is linked to the model registry entry for the models trained on it and is retained for as long as those models remain in use or under a retention obligation.
Rationale
Provenance is what makes a dispute survivable. A dataset with no recorded origin cannot be removed from a training set when a licence is challenged or a subject exercises a right, because nobody can say which models it reached. It also underpins bias auditing, since a fairness result is only interpretable against the population the data came from. The copyright compliance policy that a general-purpose AI model provider owes, including the reservation of rights under a text and data mining opt-out, is a duty of that role and belongs to the GPAI provider bundle rather than here. AIG-010 holds the model registry this record links to and AIG-012 the quality and bias properties of the dataset itself. The public summary of training content a general-purpose model provider publishes is AIG-048 and is derived from the provenance records here; the crawl-side conduct that decides what enters the datasets is AIG-049.
Applicability (9 profiles)
Obtain the provider's provenance statement for the training data (AIG-015). Provenance of data the deployer fine-tunes or evaluates on is its own.
Art.10(2) makes the collection processes and the origin of each dataset content of the documented data governance for a high-risk system, so the provenance record is read as part of the Annex IV technical documentation and is retained with it for the Art.18(1) ten years after the system is placed on the market, which is longer than the control's own floor of as long as the models are in use.
Obtain the provider's provenance statement for the training data (AIG-015). Provenance of data the deployer fine-tunes or evaluates on is its own.
Framework Mappings (29)
| DSP-20 | Data Provenance and Transparency | partial |
| EU-AI-Art.10.2 | Data Governance — Data Preparation and Bias Management | full |
| GDPR-Art.14.3 | Timing of Indirect Collection Notice | informative |
| GDPR-Art.6.1 | Lawfulness of Processing: the Six Legal Bases | informative |
| COP-C-1.2 | Reproduce and extract only lawfully accessible copyright-protected content when crawling the World Wide Web | informative |
| COP-C-1.3 | Identify and comply with rights reservations when crawling the World Wide Web | informative |
| A.7.5 | Data provenance | full |
| AML.M0023 | AI Bill of Materials | informative |
| AML.M0025 | Maintain AI Dataset Provenance | full |
| SR-4 | Provenance | informative |
| GV-1.2-001 | Trustworthy AI Characteristics Integration | GV-1.2-001 | partial |
| GV-1.6-003 | AI System Inventory | GV-1.6-003 | informative |
| GV-6.1-008 | Third-Party AI Risk Policies | GV-6.1-008 | partial |
| MG-2.2-002 | Deployed AI System Value Maintenance | MG-2.2-002 | full |
| MG-3.1-004 | Third-Party AI Risk Monitoring and Controls | MG-3.1-004 | partial |
| MG-3.2-003 | Pre-Trained Model Monitoring | MG-3.2-003 | informative |
| MG-4.1-006 | Post-Deployment AI System Monitoring | MG-4.1-006 | partial |
| MP-2.1-001 | AI System Task and Method Definition | MP-2.1-001 | full |
| MP-2.1-002 | AI System Task and Method Definition | MP-2.1-002 | informative |
| MP-4.1-006 | AI Technology and Legal Risk Mapping | MP-4.1-006 | partial |
| MP-4.1-010 | AI Technology and Legal Risk Mapping | MP-4.1-010 | full |
| MS-1.1-001 | AI Risk Measurement Approach Selection | MS-1.1-001 | informative |
| MS-2.11-005 | AI Fairness and Bias Evaluation | MS-2.11-005 | partial |
| MS-2.5-005 | AI System Validity and Reliability | MS-2.5-005 | full |
| MS-2.6-002 | AI System Safety Risk Evaluation | MS-2.6-002 | informative |
| MS-2.9-002 | AI Model Explainability and Validation | MS-2.9-002 | informative |
| GOVERN 6.1 | Third-Party AI Risk Policies | partial |
| LLM04 | Supply Chain | informative |
| LLM05 | Data and Model Poisoning | informative |
Evidence (3)
Training data provenance record for each production model, documenting origin (internal, third-party, web-scraped, synthetic), licence and copyright status, collection date, transformations applied, and legal basis for use.
Example: Data Provenance Record · LLM Fine-Tune Dataset v2 (Confluence), listing 4 source datasets: internal CRM exports (contract basis), licensed Common Crawl subset (licence agreement #CC-2024-07), synthetic augmentation (internal generation), with copyright review completed by legal 2025-05-10
Test: Request the provenance records for a sample of datasets used by production models. Verify: (1) the origin of each source is recorded, (2) licence and copyright status is recorded for each source, (3) the lawful basis relied on for any personal data is stated, (4) the transformations applied are described, (5) the record is linked to the model registry entry for the models trained on it, (6) records for retired model versions are retained while those versions remain in use or under a retention obligation.
Licence agreements or data processing agreements for third-party or licensed training datasets, confirming the organisation has the legal right to use the data for AI training purposes.
Example: Data Licence Agreement with DataProvider Ltd (executed 2024-07-15), explicitly granting rights to use dataset for model training, specifying permitted use scope and restrictions on redistribution of derivative models
Test: Request licence or data processing agreements for all third-party training datasets identified in provenance records. Verify: (1) the agreement explicitly permits use for AI/ML model training, (2) any restrictions on derivative models are identified and assessed against current use, (3) agreements are current (not expired), (4) agreements are stored in a retrievable contract repository.
Dataset lineage export from the data catalogue linking each production model to the datasets it was trained on and the recorded source of each.
Example: Data catalogue lineage export, 2026-08-31: 9 production models, 14 datasets, each carrying source, acquisition route, licence reference and version
Test: Export the dataset lineage for every model marked production in the model registry. Verify: (1) each production model resolves to at least one dataset, (2) each dataset carries a recorded source and acquisition route, (3) each dataset carries the version the model was trained against rather than the current version alone, (4) a dataset with no recorded source appears as a gap in the export rather than being absent from it, (5) the datasets in the export reconcile with the provenance records held for the same models.
Questions (2)
Is the provenance of every dataset used to train, fine-tune or evaluate a production model recorded?
Answer for the datasets behind the models currently in production. A record that covers the most recent dataset but not the ones earlier versions were trained on does not meet the control, because the obligation attaches to the model for as long as it is in use.
What does your training data provenance record include for each data source?
All six elements are expected for any source used by a production model. The link to the model registry is the element most often absent and the one that decides whether a disputed source can be traced to the models that consumed it.