GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-018 AI System Operational Monitoring

Tier 2+AIProviderDeployerGPAI Model ProviderManaged Service Provider

Description

Each production AI system has a monitoring plan naming the metrics tracked, the alert threshold for each, the review interval and the person accountable for reviewing alerts. The metrics include AI-specific measures such as output confidence distribution, refusal or null rate and output category distribution, alongside error rate and latency. Monitoring is active from the system's first production request. Monitoring outputs are reviewed at the interval the plan defines and each alert carries a triage record showing the investigation and its outcome. The plan also states how performance data reported by deployers and drawn from other sources is collected and analysed across the system's service life, including its interaction with other AI systems where it is deployed alongside them, and how that data is used to evaluate whether the system still meets the requirements it was placed on the market against. The plan is held with the system's technical documentation and versioned with it. Where a production AI system runs agents that select their own actions, the plan names each agent's authorised objective and the indicators of a departure from it, among them a significant change to the planned objective or the task hierarchy, a new objective unrelated to the assigned one, a tool call inconsistent with that objective, an attempt to reach a resource outside the authorised scope and a planning sequence that widens the authority the agent needs. A departure raises an alert to the person the plan names accountable, and its triage record states the action taken on the run, which is a pause, a restriction of the agent's tool access, a referral for approval before the run continues, a return to the authorised plan or a termination.

Rationale

An AI system degrades in ways infrastructure metrics cannot see: latency and error rate stay flat while the confidence distribution shifts and the refusal rate climbs. Generic application performance tooling will report a healthy service throughout. The review interval belongs to the tier model rather than to this text, because the same monitoring plan is read at different depths depending on the assurance required. AIG-019 holds the scheduled performance evaluation against a benchmark and this control the operational plan and the alert review; the event record the monitoring draws on is AIG-020. The post-market half answers a further question: not whether the system is healthy now, but whether what the provider claimed about it at release still holds after a year of real inputs. Deployer-reported data is the only route to that answer for a system the provider does not operate. The scope-drift limb reads the agent rather than the model. AIG-019 evaluates performance against a benchmark, AIG-031 detects a user misusing the service from outside and AIG-042 sets the permission set an out-of-scope call is read against; none of them asks whether what the agent is now trying to do is still the task it was given. AIG-023 holds the override an authorised operator invokes, where the action recorded here is the response the monitoring plan defines.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredrole duty

Monitoring in operation and informing the provider of risks are the deployer's (Art.26.4). The post-market monitoring plan and the analysis of data reported by deployers (Art.72) are the provider's.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.72 makes post-market monitoring a documented system proportionate to the technology and the risk, actively and systematically collecting and analysing performance data supplied by deployers or drawn from other sources across the system's whole lifetime. It fixes the question the data answers: whether the system still complies with the Chapter III Section 2 requirements, not only whether it still performs. Where relevant the analysis reaches the interaction with other AI systems. Art.72(3) makes the plan part of the Annex IV technical documentation, so it is an artefact produced to an authority alongside AIG-015 rather than an internal runbook. A Commission template for it is due by 2 September 2027 (ADR-025).

Public Body Deployer (EU)stablerequiredrole duty

Art.26(5) binds both seats. The deployer monitors operation against the instructions for use and informs the provider of the risks it identifies; where a risk to health, safety or fundamental rights appears it notifies the provider and the market surveillance authority immediately and suspends use of the system. Suspension is the part a monitoring plan usually lacks, so the plan names the person who can stop the system and the indicator that triggers it. AIG-021 carries the incident and reporting limb of the same paragraph.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (22)

MDS-10Model Continuous Monitoringfull
EU-AI-Art.26.4Deployer Obligations — Operational Monitoring and Incident Notificationfull
EU-AI-Art.72Post-Market Monitoring — System and Documented Planfull
COP-S-3.5Post-market monitoringinformative
A.6.2.6AI system operation and monitoringfull
AML.M0038AI Agent Scope Drift Detectionfull
GV-6.2-004Third-Party Failure Contingency Processes | GV-6.2-004full
MG-2.2-003Deployed AI System Value Maintenance | MG-2.2-003informative
MG-3.2-006Pre-Trained Model Monitoring | MG-3.2-006partial
MG-4.1-002Post-Deployment AI System Monitoring | MG-4.1-002full
MG-4.1-004Post-Deployment AI System Monitoring | MG-4.1-004informative
MG-4.1-007Post-Deployment AI System Monitoring | MG-4.1-007informative
MG-4.2-001Continual Improvement Integration | MG-4.2-001partial
MP-2.2-002AI System Knowledge Limits Documentation | MP-2.2-002informative
MP-5.2-001External Impact Feedback Practices | MP-5.2-001partial
MS-2.6-005AI System Safety Risk Evaluation | MS-2.6-005partial
MS-4.2-002Trustworthiness Measurement with Expert Input | MS-4.2-002informative
MANAGE 4.1Post-Deployment AI System Monitoringfull
MEASURE 2.4AI System Production Monitoringfull
MEASURE 3.1AI Risk Identification and Trackingpartial
ASI01Agent Goal Hijackinformative
ASI10Rogue Agentsinformative

Evidence (3)

recorddocumentmanual

Monitoring plan or monitoring configuration for each production AI system, specifying tracked metrics, alert thresholds, monitoring cadence, and named monitoring owner.

Example: Datadog monitor configuration export for ai-fraud-detection service: monitors for inference error rate (alert >2%), p95 latency (alert >800ms), null/refusal rate (alert >5%), output confidence distribution (alert if mean <0.7), owner tag: ml-ops-team, cadence: real-time streaming with daily digest review

Test: Request monitoring configuration or plan for a sample of production AI systems. Verify: (1) monitored metrics include AI-specific measures (confidence distribution, refusal/null rate, output category distribution) in addition to infrastructure metrics, (2) alert thresholds are defined for each metric, (3) a named owner is assigned, (4) a triage process for alerts is documented and accessible, (5) monitoring was active from system go-live (check monitor creation date vs deployment date). (6) where the system runs agents that select their own actions, the plan names each agent's authorised objective, the indicators of a departure from it and the action the run takes on a departure.

logtechnicalautomated

AI system monitoring review records (alert history and response logs) demonstrating that alerts are reviewed at the defined frequency and trigger a documented triage response.

Example: Datadog incident log for ai-recommendation-engine (last 90 days): 3 alerts triggered, each with a linked incident record in PagerDuty showing triage start time, investigation notes, and resolution action

Test: Request monitoring review records for a 90-day sample period. Verify: (1) alerts were reviewed within the SLA the monitoring plan defines, (2) each alert has a corresponding triage record, (3) the review cadence matches the cadence the plan sets, (4) no alert was closed without an investigation record. (5) where agents select their own actions, the alert history includes departures from the authorised objective, and the triage record for each names the action taken on the run rather than closing on a note.

recorddocumentmanual

Post-market monitoring plan for a production AI system, held with its technical documentation, stating the field data intake, the analysis cadence and the conformity evaluation the data feeds.

Example: Post-market monitoring plan, classification-api v4, rev 3 dated 2026-07-04

Test: Select a production AI system and request its post-market monitoring plan. Verify: (1) the plan names the routes by which deployer and user performance data reaches the provider and the owner of each route, (2) field data from the last two review periods is present and analysed rather than only collected, (3) the analysis records a conclusion on whether the system still meets the requirements it was placed on the market against, with the evidence for that conclusion, (4) where the system runs alongside other AI systems, the interaction between them is analysed, (5) the plan's version matches the technical documentation version for the deployed release.

Questions (3)

boolean

Does each production AI system have a documented monitoring plan?

AI system behaviour degrades in ways not visible from infrastructure metrics alone. Monitoring must include AI-specific measures: confidence score distributions, null or refusal rates, output category distributions, in addition to standard latency and error rate metrics.

multi

Which AI-specific metrics are included in your production monitoring for AI systems?

Output confidence score distributionNull rate or refusal rateOutput category or label distributionHuman override or escalation rateInput data distribution shiftsModel error rate (distinct from application error rate)Departure of an agent from its authorised objective, where the system runs agents that select their own actionsNone of the above

Programmes that monitor only latency and error rates are using generic application performance tooling, which misses the behavioural degradation patterns specific to AI systems. The last item applies only where agents select their own actions; where they do, it is the metric that shows an agent pursuing something other than the task it was given.

multi

Which of the following does the monitoring plan record for each production AI system?

An alert threshold for each metric trackedThe interval at which monitoring output is reviewedThe person accountable for reviewing alertsA named route for deployers and users to report performance dataAnalysis of that field data at a defined cadenceAn evaluation of whether the system still meets the requirements it was released againstAnalysis of the system's interaction with other AI systems it runs alongsideThe plan is held and versioned with the system's technical documentationNone of the above

Options run from the most commonly in place to the least. Operational monitoring answers whether the system is healthy now; the field data limb answers whether what was claimed about it at release still holds. The conformity evaluation is the element most often missing, because collecting field data is easier than drawing a conclusion from it.