AIG-026 AI Security and Adversarial Robustness
Description
AI systems are evaluated for AI-specific security vulnerabilities as part of verification and validation and at defined intervals in production. The evaluation covers data poisoning, model poisoning, adversarial examples, model inversion and extraction attacks and confidentiality attacks. For models that may have memorised training data the evaluation includes membership inference testing, direct extraction probing for known sensitive training data and a review of system prompt configurations that could aid extraction, scored against a documented threshold; a system above the threshold carries a recorded mitigation such as output filtering, constrained output length or differential privacy in training and is re-tested after the mitigation. Security findings are tracked to remediation. AI security evaluation is distinct from and supplementary to application security testing.
Rationale
Poisoning, adversarial examples and extraction are not found by static or dynamic application security tooling, which reads code and traffic and not model behaviour, so the evaluation has to be scoped and run separately. The extraction limb exists because a model that has memorised training data leaks it through ordinary use rather than through a vulnerability. The standardised-protocol adversarial testing owed by a general-purpose AI model provider with a systemic risk designation is a duty of that role, as is the protection of model weights and serving infrastructure that goes with it. Both belong to the GPAI provider bundle and not here. AIG-029 holds prompt injection, AIG-025 fairness evaluation and AIG-008 the verification and validation gate. Evaluation of a designated general-purpose model against standardised protocols set outside the organisation is AIG-051; protection of its weights and infrastructure to a stated security goal is AIG-052. Poisoning of the durable memory an agent writes at run time is AIG-056; this control evaluates the training and model surface.
Applicability (9 profiles)
Obtain the provider's AI security evaluation summary (AIG-034). Testing of what the deployer builds on top is APP-005 and AIG-029.
Art.15(5) names the attack classes the evaluation has to reach, data poisoning, model poisoning, adversarial examples, confidentiality attacks and model flaws. It states the standard as the system's resilience against attempts by unauthorised third parties to alter its use, outputs or performance. The measured thing is the resilience and not the existence of a test, so a result that records only that an evaluation ran does not answer it.
Obtain the provider's AI security evaluation summary (AIG-034). Testing of what the deployer builds on top is APP-005 and AIG-029.
Framework Mappings (52)
| DSP-21 | Data Poisoning Prevention & Detection | partial |
| DSP-22 | Privacy Enhancing Technologies | informative |
| MDS-06 | Adversarial Attack Analysis | full |
| MDS-07 | Robustness against Adversarial Attack / Model Hardening | partial |
| TVM-04 | Threat Analysis and Modelling | informative |
| EU-AI-Art.15.3 | Accuracy, Robustness and Cybersecurity — Cybersecurity Against AI-Specific Attacks | full |
| GDPR-Art.25.1 | Data Protection by Design | informative |
| GDPR-Art.32.1 | Technical and Organisational Security Measures | informative |
| COP-C-1.4 | Mitigate the risk of copyright-infringing outputs | informative |
| COP-S-3.2 | Model evaluations | informative |
| COP-S-5 | Safety mitigations | informative |
| 8.11 | Data masking | informative |
| AML.M0002 | Predictive AI Output Obfuscation | informative |
| AML.M0003 | Predictive AI Model Hardening | informative |
| AML.M0006 | Predictive AI Ensembles | informative |
| AML.M0007 | Sanitize Training Data | informative |
| AML.M0008 | Validate AI Model | informative |
| AML.M0010 | Predictive AI Input Restoration | informative |
| AML.M0015 | Predictive AI Adversarial Input Detection | informative |
| AML.M0035 | AI Red Team | partial |
| GV-1.2-002 | Trustworthy AI Characteristics Integration | GV-1.2-002 | informative |
| GV-4.1-002 | Safety-First Organisational Culture | GV-4.1-002 | informative |
| MG-1.3-002 | High-Priority Risk Response Planning | MG-1.3-002 | informative |
| MG-2.2-009 | Deployed AI System Value Maintenance | MG-2.2-009 | informative |
| MG-3.1-002 | Third-Party AI Risk Monitoring and Controls | MG-3.1-002 | informative |
| MP-2.3-005 | Scientific Integrity and Testing Considerations | MP-2.3-005 | full |
| MP-4.1-001 | AI Technology and Legal Risk Mapping | MP-4.1-001 | partial |
| MP-4.1-009 | AI Technology and Legal Risk Mapping | MP-4.1-009 | partial |
| MP-5.1-005 | Impact Likelihood and Magnitude Documentation | MP-5.1-005 | full |
| MP-5.1-006 | Impact Likelihood and Magnitude Documentation | MP-5.1-006 | informative |
| MS-1.1-008 | AI Risk Measurement Approach Selection | MS-1.1-008 | partial |
| MS-2.10-001 | AI Privacy Risk Examination | MS-2.10-001 | full |
| MS-2.2-004 | Human Subject Evaluation Requirements | MS-2.2-004 | informative |
| MS-2.5-006 | AI System Validity and Reliability | MS-2.5-006 | partial |
| MS-2.6-007 | AI System Safety Risk Evaluation | MS-2.6-007 | partial |
| MS-2.7-001 | AI System Security and Resilience Evaluation | MS-2.7-001 | partial |
| MS-2.7-002 | AI System Security and Resilience Evaluation | MS-2.7-002 | partial |
| MS-2.7-004 | AI System Security and Resilience Evaluation | MS-2.7-004 | informative |
| MS-2.7-007 | AI System Security and Resilience Evaluation | MS-2.7-007 | partial |
| MS-2.7-008 | AI System Security and Resilience Evaluation | MS-2.7-008 | informative |
| MS-2.7-009 | AI System Security and Resilience Evaluation | MS-2.7-009 | full |
| MS-2.8-002 | AI Transparency and Accountability Risks | MS-2.8-002 | informative |
| MS-4.2-001 | Trustworthiness Measurement with Expert Input | MS-4.2-001 | partial |
| MEASURE 2.10 | AI Privacy Risk Examination | partial |
| MEASURE 2.6 | AI System Safety Risk Evaluation | partial |
| MEASURE 2.7 | AI System Security and Resilience Evaluation | full |
| ASI06 | Memory & Context Poisoning | informative |
| LLM01 | Prompt Injection | informative |
| LLM02 | Sensitive Information Disclosure | partial |
| LLM05 | Data and Model Poisoning | full |
| LLM06 | Unbounded Consumption | informative |
| LLM08 | Hidden Context Exposure | informative |
Evidence (4)
AI-specific security evaluation report covering data poisoning, adversarial examples, model inversion, and extraction attack vectors, produced before deployment and periodically in production.
Example: AI Security Assessment · Recommendation Engine v2 (Confluence, dated 2025-10-12): adversarial robustness testing results (FGSM, PGD attacks), model inversion test result, data poisoning resilience test, findings tracked in security backlog AI-SEC-2025-003 through AI-SEC-2025-007
Test: Request the AI security evaluation reports for a sample of production AI systems. Verify: (1) the evaluation covers the AI-specific attack classes the control names and is distinct from static and dynamic application security results, (2) the poisoning limb covers ATLAS AML.T0020 Training Data Poisoning, AML.T0018.000 Poison AI Model, AML.T0059 Erode Dataset Integrity and AML.T0043.004 Insert Backdoor Trigger, and the confidentiality limb covers AML.T0024 Exfiltration via AI Inference API and AML.T0057 LLM Data Leakage, (3) findings are tracked to remediation with a severity and a status, (4) an evaluation exists dated before deployment and a further evaluation within the defined interval, (5) the attack classes evaluated match the classes the system's threat model identifies, and every technique named above is either covered or recorded as out of scope with a stated reason.
Third-party red team assessment or AI security audit report for high-risk AI systems, providing independent validation of AI-specific security posture.
Example: Red Team Assessment Report · Customer Support Assistant v4 (SecurityVendor Ltd, 2025-09-15), covering adversarial robustness, jailbreak resilience, model extraction, membership inference and training data leakage, with risk ratings and remediation recommendations
Test: Request the most recent independent AI security assessment. Verify: (1) the scope covers AI-specific attack vectors beyond conventional application security, (2) the assessor is independent of the team that built the system, (3) findings are tracked and remediation is evidenced, (4) the report falls within the interval defined for the system, (5) the scope covers the attack classes the system's threat model identifies.
Training data extraction risk evaluation report covering membership inference testing, direct extraction probing for sensitive training data, and system prompt configuration review, produced before deployment and after retraining.
Example: Training Data Extraction Risk Report · LLM Customer Support Model v2 (Confluence, 2025-11-20): membership inference test (500 member/non-member pairs, AUC 0.53, near random, PASS), extraction probing (200 prompts targeting known training data patterns, 0 verbatim extractions > 50 tokens), system prompt injection review: PASS; risk score: LOW; no additional mitigations required
Test: Request the training data extraction risk evaluation report for each production LLM. Verify: (1) membership inference testing was performed with a documented methodology and the AUC result is recorded and compared to a defined threshold, covering ATLAS AML.T0024.000 Infer Training Data Membership, (2) direct extraction probing was performed for known sensitive training data categories (PII, credentials, copyrighted content), covering ATLAS AML.T0057 LLM Data Leakage, with AML.T0024.001 Invert AI Model and AML.T0024.002 Extract AI Model either covered or recorded as out of scope with a stated reason, (3) system prompt configuration was reviewed for extraction facilitation risks, covering ATLAS AML.T0056 Extract LLM System Prompt, (4) report is dated before initial deployment and after any retraining event, (5) where risk score exceeds the documented threshold, at least one mitigation (output filtering, differential privacy, constrained output length) is implemented and evidenced.
Output filtering or differential privacy configuration applied to LLMs scoring above the training data extraction risk threshold, demonstrating that mitigations are technically enforced.
Example: LLM output filtering configuration (AWS Bedrock Guardrails or custom post-processing pipeline): max_verbatim_output_tokens: 150, PII_redaction: enabled (regex + NER model), known_credential_pattern_filter: enabled, exact_match_training_data_filter: enabled (hash-based bloom filter against known sensitive training records), filter_version: v2 (git tag output-filter-v2)
Test: For any LLM that scored above the extraction risk threshold in its evaluation report, request the mitigation configuration. Verify: (1) output length constraint is configured and enforced (test with a verbatim reproduction prompt), (2) PII redaction is active and covers the PII categories present in the training data, (3) known credential or sensitive pattern filters are configured, (4) configuration is version-controlled and the review date postdates the most recent extraction risk evaluation, (5) mitigation effectiveness is validated in the evaluation report (post-mitigation re-test result is present).
Questions (3)
Are AI systems evaluated for AI-specific security vulnerabilities?
AI-specific attacks are not detected by conventional SAST/DAST tooling. Evaluation must be explicitly scoped to AI attack classes and conducted separately from standard application security testing.
Which AI-specific attack classes are evaluated in your AI security testing programme?
Coverage of all six attack classes characterises a mature AI security evaluation. The results belong in a report of their own: an application penetration test that mentions the AI endpoint does not evidence this control, because it tests the interface rather than the model.
Which of the following training data extraction controls are applied to your LLM-based systems?
Membership inference testing and extraction probing are the minimum baseline. Mitigations (output length constraints, PII redaction, differential privacy) are required for any LLM that scores above the documented risk threshold in its evaluation. The risk evaluation methodology and results must be retained.