GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-051 Model Evaluation Under Standardised Protocols

Tier 3+AIGeneral-purpose modelgpai-systemic

Description

Each general-purpose AI model designated as carrying systemic risk is evaluated against the standardised protocols and state-of-the-art tools the applicable code of practice and the AI Office name, before it is placed on the market and at each trigger point the systemic risk framework sets. The evaluation covers each identified systemic risk, includes documented adversarial testing by evaluators independent of the team that trained the model, with the elicitation methods, the compute and the model access they had recorded, and for chemical, biological, radiological, nuclear and offensive cyber capabilities runs to a test plan and response policy settled before development began. Each evaluation record carries the results against the acceptance criteria, the mitigations adopted for the risks found and the re-evaluation after each mitigation.

Rationale

The point of standardised protocols is comparability: the AI Office reads evaluations across providers, so an in-house benchmark that cannot be compared does not pass. Recording the elicitation effort, compute and access the evaluators had is what lets a reader judge whether a null result means the capability is absent or was not found; the Code's safety margin for under-elicitation rests on that record. The pre-development test plan for CBRN and offensive cyber capability is the one element that has to exist before the training run rather than after it. Boundary with AIG-026: that control tests the systems the organisation operates on its own protocol; this one tests a designated model on protocols set outside the organisation. Boundary with AIG-008: that control is the deployment gate for a system; the evaluation here feeds the AIG-050 determination for a model. gpai-provider seat (ADR-031).

Applicability (9 profiles)

SaaS AI Providerstablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

Enterprise AI Deployerstablenot-applicableout of scope

A general-purpose model providers duty (gpai-model-provider, ADR-046). The deployer takes the model documentation and the published training summary the provider issues into its AIG-032 assessment.

GPAI Model Providerstableconditionalrisk class duty

Condition: ai_risk_class in gpai-systemic

Art.55(1)(a) binds the provider of a model designated as carrying systemic risk.

High-Risk Provider (EU)stablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

Public Body Deployer (EU)stablenot-applicableout of scope

A general-purpose model providers duty (gpai-model-provider, ADR-046). The deployer takes the model documentation and the published training summary the provider issues into its AIG-032 assessment.

Data Act Cloud Provider (EU)stablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

DORA ICT Provider (EU)stablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

HIPAA Business Associate (US)stablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

NIS2 Cloud Provider (EU)stablenot-applicableout of scope

Evaluation under standardised protocols attaches to a designated general-purpose model (ADR-046). A SaaS provider evaluates the systems it operates under AIG-026 and AIG-008.

Framework Mappings (6)

MDS-06Adversarial Attack Analysisinformative
EU-AI-Art.55.1Systemic Risk Obligations — Adversarial Testing and Model Evaluationfull
COP-S-3.2Model evaluationsfull
COP-S-5Safety mitigationspartial
AML.M0035AI Red Teaminformative
GV-1.3-003Risk Management Activity Level Determination | GV-1.3-003full

Evidence (3)

reportdocumentmanual

The evaluation report for a release or a trigger, covering each systemic risk against the named protocol, with the adversarial testing record and the results against the acceptance criteria.

Example: Aurora-2 v2.3 systemic risk evaluation report EV-2026-09, 14 August 2026, three external evaluators, protocols listed at section 2.

Test: Verify: (1) the report names the protocol and tools used for each systemic risk and they are ones the code or the AI Office names, (2) the evaluators are recorded as independent of the training team, with their elicitation methods, compute budget and model access, (3) each result is stated against the acceptance criteria of the framework rather than as a raw score, (4) risks found above a criterion carry the mitigation adopted and the re-evaluation result after it, (5) the evaluation date precedes the placement date or falls within the period the framework allows after a trigger.

policydocumentmanual

The test plan and response policy for chemical, biological, radiological, nuclear and offensive cyber capabilities, dated before the training run it governs.

Example: Dangerous Capability Test Plan and Response Policy v2, approved 3 February 2026, governing the Aurora-2 training run that began 17 March 2026.

Test: Verify: (1) the plan is dated before the start of the training run it governs, shown against the run record in the model registry, (2) it names the capabilities evaluated and the evaluation points across development, (3) the response policy states what follows a positive result, including pausing the run or withholding release, (4) each evaluation point in the plan has a corresponding result in the evaluation records, (5) a positive result in the period shows the response the policy prescribes.

tool_outputtechnicalautomated

Raw output of the evaluation harness for a named protocol, reconciled to the figures the evaluation report states.

Example: Evaluation harness run 2026-08-12-03 for the cyber capability protocol, 1,240 tasks, artefacts in the evaluation store.

Test: Verify: (1) the harness output exists for each protocol the report names, (2) the scores in the report reproduce from the output within the tolerance the protocol states, (3) the model version evaluated matches the release under determination, (4) the elicitation settings recorded in the output match those the report describes.

Questions (3)

boolean

Is each designated model evaluated against standardised protocols before placement and at each framework trigger?

Standardised means protocols the code of practice or the AI Office names, so results can be compared across providers. Evaluation on an internal benchmark alone is a no.

multi

Which of the following does the evaluation programme include?

Every identified systemic risk is coveredAdversarial testing by evaluators independent of the training teamThe evaluators' elicitation methods, compute and model access are recordedA test plan and response policy for CBRN and offensive cyber capability settled before developmentResults stated against the acceptance criteriaMitigations adopted for risks found and re-evaluation after eachNone of the above

Options follow the evaluation from scope to remediation. The pre-development test plan is the item that cannot be recovered afterwards; if it post-dates the training run, leave it unticked.

select

Who performs the adversarial testing?

External evaluators with recorded access and resourcesAn internal team independent of the training teamThe training team itselfNo adversarial testing is performed

Options run from the strongest independence to none. Where both external and internal evaluators are used, answer for the external.