GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-031 AI Misuse, Jailbreak and Abuse Detection

Tier 2+AIGenerativeProviderDeployerGPAI Model ProviderManaged Service Provider

Description

AI systems exposed to external users have mechanisms to detect and respond to misuse attempts. Detection covers jailbreak patterns, requests for prohibited content categories, automated abuse and abnormal usage patterns by volume, sequence or content. Detected violations trigger a documented response: rate limiting, session termination or account action. Misuse patterns are reviewed periodically to update detection logic, and AI-specific rate limits are documented and applied independently of generic API rate limiting. For each generative model or modality the organisation places on the market, a recorded assessment states which prohibited generation outcomes are reasonably foreseeable and reproducible without significant technical modification, and names the technical safeguard that prevents each of them rather than detects it after the fact. Misuse that is observed or reported feeds a correction of those safeguards, and the correction is recorded against the assessment. A resource ceiling applies to each request, each inference job and each agent run, bounding input size, batch size, execution time, memory, compute and output size, and for a generative service the context and output tokens. An agent run is bounded on iterations, retries, tool calls, parallel tasks, delegation depth and downstream spend. A budget ceiling applies per credential, per user and per account. The bounds apply across the whole workflow one request starts, and reaching a bound halts the work rather than raising an alert, with the halt recorded.

Rationale

Public-facing generative systems face continuous adversarial probing, and misuse detection is an AI-specific operational security control with no equivalent in classical application security. The foreseeability assessment answers a different question from the detection stack: not whether misuse is being caught, but whether the system's design, training and user-facing functionality make a prohibited outcome reachable without real effort, and whether anything stops it before it happens. Where the law turns on foreseeability and on the adequacy of safeguards, an undocumented judgement is not a defence, and a detection-only posture concedes the point. A rate limit governs how many requests arrive; a ceiling governs what one of them may consume, which is why a single recursive agent run can exhaust a budget that no rate limit was ever breached to reach. INF-012 monitors utilisation against thresholds and alerts before exhaustion affects availability, and halts nothing: the halt is stated here. A forged audio, image or video artefact submitted as evidence of identity, provenance or an event is AIG-057; this control detects misuse of the service by a user.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredcore

Runtime detection and response on systems the deployer exposes to users are its own. The assessment of prohibited generation outcomes for each model placed on the market is the provider's.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.9(2) requires the risk management system to estimate risk under reasonably foreseeable misuse as well as intended use, which is the question the foreseeability assessment in this control answers. It makes that assessment an input to Art.9 rather than a safety-team artefact. Art.15(5) covers the misuse that alters the system's use, outputs or performance, so the detection and response stack is measured as resilience. The Art.5(1)(ba) and (bb) safeguards are a prohibition binding whatever the risk class and apply from 2 December 2026 (ADR-025), ahead of the Chapter III dates.

Public Body Deployer (EU)stablerequiredcore

Runtime detection and response on systems the deployer exposes to users are its own. The assessment of prohibited generation outcomes for each model placed on the market is the provider's.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (37)

AIS-09Input Validationpartial
AIS-10Output Validationinformative
GRC-09Acceptable Use of the AI Serviceinformative
LOG-15Input Monitoringinformative
LOG-16Output Monitoringinformative
TVM-13Guardrailsinformative
EU-AI-Art.15.3Accuracy, Robustness and Cybersecurity — Cybersecurity Against AI-Specific Attackspartial
EU-AI-Art.5.1iProhibited — Non-Consensual Intimate or Sexually Explicit Materialpartial
EU-AI-Art.5.1jProhibited — Child Sexual Abuse Materialpartial
EU-AI-Art.5.1kProhibited Practices — Scope Conditions for Intimate and Child Sexual Abuse Materialfull
EU-AI-Art.9.2AI Risk Management System — Risk Identification and Analysispartial
COP-S-5.1Appropriate safety mitigationsinformative
AML.M0004Limit AI Service Query Volume and Ratefull
AML.M0019Control Access to AI Models and Data in Productioninformative
AML.M0020Generative AI Guardrailsinformative
AML.M0022Generative AI Model Alignmentinformative
AML.M0036Limit AI Workload Resource Consumptionfull
AML.M0038AI Agent Scope Drift Detectioninformative
GV-1.4-001Transparent Risk Management Policies | GV-1.4-001full
GV-3.2-003Human-AI Configuration Roles | GV-3.2-003informative
MG-2.2-005Deployed AI System Value Maintenance | MG-2.2-005full
MG-3.1-004Third-Party AI Risk Monitoring and Controls | MG-3.1-004informative
MG-3.2-004Pre-Trained Model Monitoring | MG-3.2-004informative
MG-3.2-005Pre-Trained Model Monitoring | MG-3.2-005full
MP-1.1-003AI System Purpose and Deployment Context | MP-1.1-003informative
MP-1.1-004AI System Purpose and Deployment Context | MP-1.1-004full
MP-5.1-001Impact Likelihood and Magnitude Documentation | MP-5.1-001informative
MS-2.6-006AI System Safety Risk Evaluation | MS-2.6-006full
MS-2.6-007AI System Safety Risk Evaluation | MS-2.6-007informative
MS-2.7-007AI System Security and Resilience Evaluation | MS-2.7-007informative
MS-2.8-001AI Transparency and Accountability Risks | MS-2.8-001partial
MANAGE 4.1Post-Deployment AI System Monitoringinformative
MEASURE 3.3User and Community Feedback Processesinformative
ASI01Agent Goal Hijackinformative
ASI02Tool Misuse and Exploitationinformative
ASI10Rogue Agentspartial
LLM06Unbounded Consumptionpartial

Evidence (3)

configurationtechnicalautomated

Jailbreak and misuse detection configuration for public-facing LLM systems, including detection rule set, violation response actions (rate limit, session termination, account action), and AI-specific rate limit settings. The export also carries the per-request, per-run and per-account resource and budget ceilings and the action taken when one is reached.

Example: AWS Bedrock Guardrails + API gateway rate-limit configuration export: jailbreak_detection: enabled, blocked_topic_categories: [violence, CSAM, credential_theft], rate_limit_per_user: 100_requests/hour (AI-specific, separate from API gateway default 1000/hour), violation_action: session_terminate + alert_security_ops, pattern_review_cadence: monthly

Test: Request the misuse and jailbreak detection configuration. Verify: (1) jailbreak detection is enabled with a named rule set or model, (2) prohibited content categories are enumerated in the configuration, (3) AI-specific rate limits are configured separately from generic API rate limits and are documented, (4) violation response actions are defined (not just detection/logging), (5) abnormal usage pattern detection is configured (volume, sequence anomalies), (6) misuse pattern review cadence is documented. (7) a resource ceiling is configured for a request, for an inference job and for an agent run, covering at least input size, execution time, tokens and estimated cost, and for an agent run the step count, the delegation depth and the downstream spend, (8) a budget ceiling exists per credential, per user and per account, (9) a test workload driven past a ceiling stops rather than continuing with an alert raised, and the halt appears in the record.

logtechnicalautomated

Jailbreak and policy violation detection logs covering a 90-day sample period, demonstrating that violations are detected, response actions are triggered, and patterns are reviewed.

Example: Security operations log extract (Splunk, last 90 days) for ai-customer-chatbot: 47 jailbreak attempts detected, 12 policy violations, all 59 events triggered session_terminate action within 200ms, weekly review tickets AI-SEC-2026-W01 through AI-SEC-2026-W13 showing pattern review completed

Test: Request jailbreak and violation detection logs for a 90-day period. Verify: (1) detection events are present and timestamped, (2) each detection event shows a response action was triggered (not detection-only), (3) periodic review events are recorded confirming that misuse patterns were reviewed to update detection rules, (4) bot-driven abuse events (high-volume automated patterns) are distinguishable in the log and responded to.

recorddocumentmanual

Foreseeability and safeguard assessment for each generative model or modality, listing the prohibited outcomes considered, the safeguard that prevents each and the evidence that the safeguard holds.

Example: Foreseeability assessment, image-gen v5, rev 2 dated 2026-06-27

Test: Request the assessments. Verify: (1) every generative modality the product offers has an assessment dated before that modality was released, (2) each prohibited outcome considered names the safeguard that prevents it and distinguishes prevention from detection, (3) the evidence for each safeguard is a test result rather than a design claim, (4) the safeguard test exercises ATLAS AML.T0054 LLM Jailbreak and AML.T0068 LLM Prompt Obfuscation against each prohibited outcome, and the abuse limb covers AML.T0029 Denial of AI Service and AML.T0034 Cost Harvesting with its sub-techniques AML.T0034.000 Excessive Queries, AML.T0034.001 Resource-Intensive Queries and AML.T0034.002 Agentic Resource Consumption, (5) the test attempted the outcome without significant technical modification of the system, matching the standard the assessment is judged against, (6) misuse observed or reported in the period resolves to a recorded change to a safeguard or a recorded decision that none was needed.

Questions (3)

boolean

Do your AI systems exposed to external users have a mechanism to detect misuse attempts?

Net-new control: public-facing LLM systems face continuous adversarial probing. Misuse detection is an AI-specific operational security control with no equivalent in classical application security and is not addressed operationally by any existing framework.

multi

Which misuse and abuse detection capabilities are active for your public-facing AI systems?

Jailbreak pattern detection (attempts to bypass safety instructions)Policy violation detection (requests for prohibited content categories)Abnormal usage pattern detection by volume, sequence or content, including automated abuseAI-specific rate limits configured independently of generic API rate limitsResource ceilings per request, per inference job and per agent run that halt the work when reachedA budget ceiling per credential, per user and per account that halts inference rather than raising an alertDocumented violation response actions (rate limiting, session termination, account action)Periodic misuse pattern review to update detection logicNone of the above

Options run from the most commonly in place to the least. AI-specific rate limits separate from generic API rate limits are often absent: shared limits allow targeted abuse to consume a disproportionate share of capacity before a generic control notices. A ceiling is a different instrument from a rate limit and is the item most often missing: it bounds what one request, job or run may consume and it stops the work, where an alert on a budget is outrun by a fast workload.

multi

For each generative model or modality you place on the market, which of the following are recorded?

The prohibited generation outcomes that are reasonably foreseeable without significant technical modificationThe technical safeguard that prevents each of those outcomesTest evidence that each safeguard holdsThe correction made to a safeguard after misuse was observed or reportedNone of the above

Options run from the most commonly recorded to the least. A detection-only posture answers whether misuse is being caught, not whether the outcome was reachable without real effort and whether anything stopped it first. Where the law turns on foreseeability and on the adequacy of safeguards, an undocumented judgement is not a defence.