AIG-051 Model Evaluation Under Standardised Protocols
Description
Each general-purpose AI model designated as carrying systemic risk is evaluated against the standardised protocols and state-of-the-art tools the applicable code of practice and the AI Office name, before it is placed on the market and at each trigger point the systemic risk framework sets. The evaluation covers each identified systemic risk, includes documented adversarial testing by evaluators independent of the team that trained the model, with the elicitation methods, the compute and the model access they had recorded, and for chemical, biological, radiological, nuclear and offensive cyber capabilities runs to a test plan and response policy settled before development began. Each evaluation record carries the results against the acceptance criteria, the mitigations adopted for the risks found and the re-evaluation after each mitigation.
Rationale
The point of standardised protocols is comparability: the AI Office reads evaluations across providers, so an in-house benchmark that cannot be compared does not pass. Recording the elicitation effort, compute and access the evaluators had is what lets a reader judge whether a null result means the capability is absent or was not found; the Code's safety margin for under-elicitation rests on that record. The pre-development test plan for CBRN and offensive cyber capability is the one element that has to exist before the training run rather than after it. Boundary with AIG-026: that control tests the systems the organisation operates on its own protocol; this one tests a designated model on protocols set outside the organisation. Boundary with AIG-008: that control is the deployment gate for a system; the evaluation here feeds the AIG-050 determination for a model. gpai-provider seat (ADR-031).
Applicability (9 profiles)
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
A general-purpose model providers duty (gpai-model-provider, ADR-046). The deployer takes the model documentation and the published training summary the provider issues into its AIG-032 assessment.
Condition: ai_risk_class in gpai-systemic
Art.55(1)(a) binds the provider of a model designated as carrying systemic risk.
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
A general-purpose model providers duty (gpai-model-provider, ADR-046). The deployer takes the model documentation and the published training summary the provider issues into its AIG-032 assessment.
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.
Framework Mappings (6)
| MDS-06 | Adversarial Attack Analysis | informative |
| EU-AI-Art.55.1 | Systemic Risk Obligations — Adversarial Testing and Model Evaluation | full |
| COP-S-3.2 | Model evaluations | full |
| COP-S-5 | Safety mitigations | partial |
| AML.M0035 | AI Red Team | informative |
| GV-1.3-003 | Risk Management Activity Level Determination | GV-1.3-003 | full |
Evidence (3)
The evaluation report for a release or a trigger, covering each systemic risk against the named protocol, with the adversarial testing record and the results against the acceptance criteria.
Example: Aurora-2 v2.3 systemic risk evaluation report EV-2026-09, 14 August 2026, three external evaluators, protocols listed at section 2.
Test: Verify: (1) the report names the protocol and tools used for each systemic risk and they are ones the code or the AI Office names, (2) the evaluators are recorded as independent of the training team, with their elicitation methods, compute budget and model access, (3) each result is stated against the acceptance criteria of the framework rather than as a raw score, (4) risks found above a criterion carry the mitigation adopted and the re-evaluation result after it, (5) the evaluation date precedes the placement date or falls within the period the framework allows after a trigger.
The test plan and response policy for chemical, biological, radiological, nuclear and offensive cyber capabilities, dated before the training run it governs.
Example: Dangerous Capability Test Plan and Response Policy v2, approved 3 February 2026, governing the Aurora-2 training run that began 17 March 2026.
Test: Verify: (1) the plan is dated before the start of the training run it governs, shown against the run record in the model registry, (2) it names the capabilities evaluated and the evaluation points across development, (3) the response policy states what follows a positive result, including pausing the run or withholding release, (4) each evaluation point in the plan has a corresponding result in the evaluation records, (5) a positive result in the period shows the response the policy prescribes.
Raw output of the evaluation harness for a named protocol, reconciled to the figures the evaluation report states.
Example: Evaluation harness run 2026-08-12-03 for the cyber capability protocol, 1,240 tasks, artefacts in the evaluation store.
Test: Verify: (1) the harness output exists for each protocol the report names, (2) the scores in the report reproduce from the output within the tolerance the protocol states, (3) the model version evaluated matches the release under determination, (4) the elicitation settings recorded in the output match those the report describes.
Questions (3)
Is each designated model evaluated against standardised protocols before placement and at each framework trigger?
Standardised means protocols the code of practice or the AI Office names, so results can be compared across providers. Evaluation on an internal benchmark alone is a no.
Which of the following does the evaluation programme include?
Options follow the evaluation from scope to remediation. The pre-development test plan is the item that cannot be recovered afterwards; if it post-dates the training run, leave it unticked.
Who performs the adversarial testing?
Options run from the strongest independence to none. Where both external and internal evaluators are used, answer for the external.