AIG-029 Prompt Injection Protection
Description
Prompt injection is named in the threat model of each system that places user-supplied or externally sourced content into a model prompt. Security testing for each such system exercises direct injection through user input and indirect injection through retrieved documents, tool results and other externally sourced content. The outcome of each case is recorded. Systems that call tools enforce privilege separation between the model and the tools it may invoke and run tool execution in a sandbox. Injection attempts detected at runtime are logged with the pattern that matched. Detection patterns are updated from the test results.
Rationale
Prompt injection is the attack class that has no fix, only containment: the model cannot reliably distinguish an instruction from data, so the control is testing both injection paths and limiting what a successful injection can reach. Indirect injection through a retrieved document or a tool result is the path teams miss, because it needs no hostile user. The old wording asked for any two mechanisms from a list, which a system could satisfy without ever having been tested. The authorisation model for the actions an agent may take, including approval for consequential actions, is AIG-042; the sandboxed execution of untrusted code is APP-016. AIG-026 holds the wider AI security evaluation and APP-011 the API-layer controls. Indirect injection through retrieved content is the usual route into an agent's memory; the integrity of the store once content has arrived is AIG-056.
Applicability (9 profiles)
The row is the deployer's for every integration that places its own or external content into a prompt. Where the provider's interface is used as shipped, the testing is the provider's (AIG-034 test summary).
Art.15(5) reaches prompt injection through the model flaws it names and through the alteration of the system's use and outputs it prohibits, so the direct and indirect injection tests become evidence for a Section 2 requirement rather than application security housekeeping. The mapping stays partial because Art.15(5) states the resilience and not the two injection paths, the privilege separation or the sandbox, so a conformity file that cites this control has to show the test results and not the article.
The row is the deployer's for every integration that places its own or external content into a prompt. Where the provider's interface is used as shipped, the testing is the provider's (AIG-034 test summary).
Framework Mappings (16)
| AIS-09 | Input Validation | informative |
| AIS-13 | AI Sandboxing | informative |
| AIS-15 | Prompt Differentiation | partial |
| TVM-04 | Threat Analysis and Modelling | informative |
| EU-AI-Art.15.3 | Accuracy, Robustness and Cybersecurity — Cybersecurity Against AI-Specific Attacks | partial |
| AML.M0020 | Generative AI Guardrails | informative |
| AML.M0021 | Generative AI Guidelines | informative |
| AML.M0030 | Restrict AI Agent Tool Invocation on Untrusted Data | partial |
| AML.M0033 | Input and Output Validation for AI Agent Components | informative |
| GV-3.2-005 | Human-AI Configuration Roles | GV-3.2-005 | informative |
| MP-2.3-005 | Scientific Integrity and Testing Considerations | MP-2.3-005 | informative |
| MS-2.7-007 | AI System Security and Resilience Evaluation | MS-2.7-007 | partial |
| MEASURE 2.7 | AI System Security and Resilience Evaluation | informative |
| ASI01 | Agent Goal Hijack | full |
| ASI06 | Memory & Context Poisoning | informative |
| LLM01 | Prompt Injection | full |
Evidence (2)
Serving and gateway configuration for each system in scope showing the prompt architecture separating instructions from untrusted content, the privilege separation and sandbox settings for tool execution and the runtime logging of injection attempts.
Example: LLM API gateway configuration (AWS Bedrock Guardrails export or LangChain input guard config): input_sanitisation: enabled (strips HTML/JS injection patterns), system_prompt_separation: instruction_prefix=SYSTEM_INSTRUCTION, data_prefix=USER_DATA, tool_call_sandbox: docker_isolated (no network access), injection_pattern_log: enabled, guardrail_version: v4
Test: Request the configuration for each system that places user-supplied or externally sourced content into a model prompt. Verify: (1) the prompt architecture separates instructions from untrusted content in the deployed template rather than only in a design document, (2) a tool-calling system runs tool execution in a sandbox and the model's credentials do not carry the privileges of the tools it invokes, (3) runtime detection of injection attempts is active and the log records the pattern that matched, (4) the detection configuration is version-controlled and its current version is the one deployed.
Security test report including prompt injection test cases, demonstrating that the system was evaluated for prompt injection resistance during V&V testing.
Example: Security Test Report · AI Customer Support Bot v3 (Burp Suite / custom harness, 2025-12-10): 45 prompt injection test cases executed (direct injection, indirect injection via retrieved documents, instruction override attempts), 0 successful injections, 3 partial bypasses noted and mitigated, re-test passed 2025-12-20
Test: Request the security test results covering prompt injection for each system in scope. Verify: (1) test cases exercise direct injection through user input, covering ATLAS AML.T0051 LLM Prompt Injection and its sub-technique AML.T0051.000 Direct, (2) test cases exercise indirect injection through retrieved documents, tool results or other externally sourced content, covering AML.T0051.001 Indirect, AML.T0066 Retrieval Content Crafting and AML.T0070 RAG Poisoning, (3) test cases exercise the delayed and the obfuscated variant, AML.T0051.002 Triggered and AML.T0068 LLM Prompt Obfuscation, or record each as out of scope with a stated reason, (4) the outcome of each case is recorded rather than only the summary, (5) every case that succeeded carries a remediation and a recorded re-test result, (6) the patterns discovered during testing were added to the runtime detection configuration.
Questions (2)
Is prompt injection named in the threat model of each system that places user-supplied or externally sourced content into a model prompt?
Answer for the systems in scope, not for the organisation. A threat model that names injection only as a generic input validation risk does not meet the control; the entry should identify the untrusted content sources the system consumes.
Which of the following are in place for your systems that place user-supplied or externally sourced content into a model prompt?
Every item applies to each system in scope. Indirect injection is the path most often untested: content arriving from a retrieved document, a web page or a tool response carries instructions without any hostile user being present. For a system that calls tools, the last four items are the containment that decides how far a successful injection reaches.