AIG-018 AI System Operational Monitoring
Description
Each production AI system has a monitoring plan naming the metrics tracked, the alert threshold for each, the review interval and the person accountable for reviewing alerts. The metrics include AI-specific measures such as output confidence distribution, refusal or null rate and output category distribution, alongside error rate and latency. Monitoring is active from the system's first production request. Monitoring outputs are reviewed at the interval the plan defines and each alert carries a triage record showing the investigation and its outcome. The plan also states how performance data reported by deployers and drawn from other sources is collected and analysed across the system's service life, including its interaction with other AI systems where it is deployed alongside them, and how that data is used to evaluate whether the system still meets the requirements it was placed on the market against. The plan is held with the system's technical documentation and versioned with it. Where a production AI system runs agents that select their own actions, the plan names each agent's authorised objective and the indicators of a departure from it, among them a significant change to the planned objective or the task hierarchy, a new objective unrelated to the assigned one, a tool call inconsistent with that objective, an attempt to reach a resource outside the authorised scope and a planning sequence that widens the authority the agent needs. A departure raises an alert to the person the plan names accountable, and its triage record states the action taken on the run, which is a pause, a restriction of the agent's tool access, a referral for approval before the run continues, a return to the authorised plan or a termination.
Rationale
An AI system degrades in ways infrastructure metrics cannot see: latency and error rate stay flat while the confidence distribution shifts and the refusal rate climbs. Generic application performance tooling will report a healthy service throughout. The review interval belongs to the tier model rather than to this text, because the same monitoring plan is read at different depths depending on the assurance required. AIG-019 holds the scheduled performance evaluation against a benchmark and this control the operational plan and the alert review; the event record the monitoring draws on is AIG-020. The post-market half answers a further question: not whether the system is healthy now, but whether what the provider claimed about it at release still holds after a year of real inputs. Deployer-reported data is the only route to that answer for a system the provider does not operate. The scope-drift limb reads the agent rather than the model. AIG-019 evaluates performance against a benchmark, AIG-031 detects a user misusing the service from outside and AIG-042 sets the permission set an out-of-scope call is read against; none of them asks whether what the agent is now trying to do is still the task it was given. AIG-023 holds the override an authorised operator invokes, where the action recorded here is the response the monitoring plan defines.
Applicability (9 profiles)
Monitoring in operation and informing the provider of risks are the deployer's (Art.26.4). The post-market monitoring plan and the analysis of data reported by deployers (Art.72) are the provider's.
Art.72 makes post-market monitoring a documented system proportionate to the technology and the risk, actively and systematically collecting and analysing performance data supplied by deployers or drawn from other sources across the system's whole lifetime. It fixes the question the data answers: whether the system still complies with the Chapter III Section 2 requirements, not only whether it still performs. Where relevant the analysis reaches the interaction with other AI systems. Art.72(3) makes the plan part of the Annex IV technical documentation, so it is an artefact produced to an authority alongside AIG-015 rather than an internal runbook. A Commission template for it is due by 2 September 2027 (ADR-025).
Art.26(5) binds both seats. The deployer monitors operation against the instructions for use and informs the provider of the risks it identifies; where a risk to health, safety or fundamental rights appears it notifies the provider and the market surveillance authority immediately and suspends use of the system. Suspension is the part a monitoring plan usually lacks, so the plan names the person who can stop the system and the indicator that triggers it. AIG-021 carries the incident and reporting limb of the same paragraph.
Framework Mappings (22)
| MDS-10 | Model Continuous Monitoring | full |
| EU-AI-Art.26.4 | Deployer Obligations — Operational Monitoring and Incident Notification | full |
| EU-AI-Art.72 | Post-Market Monitoring — System and Documented Plan | full |
| COP-S-3.5 | Post-market monitoring | informative |
| A.6.2.6 | AI system operation and monitoring | full |
| AML.M0038 | AI Agent Scope Drift Detection | full |
| GV-6.2-004 | Third-Party Failure Contingency Processes | GV-6.2-004 | full |
| MG-2.2-003 | Deployed AI System Value Maintenance | MG-2.2-003 | informative |
| MG-3.2-006 | Pre-Trained Model Monitoring | MG-3.2-006 | partial |
| MG-4.1-002 | Post-Deployment AI System Monitoring | MG-4.1-002 | full |
| MG-4.1-004 | Post-Deployment AI System Monitoring | MG-4.1-004 | informative |
| MG-4.1-007 | Post-Deployment AI System Monitoring | MG-4.1-007 | informative |
| MG-4.2-001 | Continual Improvement Integration | MG-4.2-001 | partial |
| MP-2.2-002 | AI System Knowledge Limits Documentation | MP-2.2-002 | informative |
| MP-5.2-001 | External Impact Feedback Practices | MP-5.2-001 | partial |
| MS-2.6-005 | AI System Safety Risk Evaluation | MS-2.6-005 | partial |
| MS-4.2-002 | Trustworthiness Measurement with Expert Input | MS-4.2-002 | informative |
| MANAGE 4.1 | Post-Deployment AI System Monitoring | full |
| MEASURE 2.4 | AI System Production Monitoring | full |
| MEASURE 3.1 | AI Risk Identification and Tracking | partial |
| ASI01 | Agent Goal Hijack | informative |
| ASI10 | Rogue Agents | informative |
Evidence (3)
Monitoring plan or monitoring configuration for each production AI system, specifying tracked metrics, alert thresholds, monitoring cadence, and named monitoring owner.
Example: Datadog monitor configuration export for ai-fraud-detection service: monitors for inference error rate (alert >2%), p95 latency (alert >800ms), null/refusal rate (alert >5%), output confidence distribution (alert if mean <0.7), owner tag: ml-ops-team, cadence: real-time streaming with daily digest review
Test: Request monitoring configuration or plan for a sample of production AI systems. Verify: (1) monitored metrics include AI-specific measures (confidence distribution, refusal/null rate, output category distribution) in addition to infrastructure metrics, (2) alert thresholds are defined for each metric, (3) a named owner is assigned, (4) a triage process for alerts is documented and accessible, (5) monitoring was active from system go-live (check monitor creation date vs deployment date). (6) where the system runs agents that select their own actions, the plan names each agent's authorised objective, the indicators of a departure from it and the action the run takes on a departure.
AI system monitoring review records (alert history and response logs) demonstrating that alerts are reviewed at the defined frequency and trigger a documented triage response.
Example: Datadog incident log for ai-recommendation-engine (last 90 days): 3 alerts triggered, each with a linked incident record in PagerDuty showing triage start time, investigation notes, and resolution action
Test: Request monitoring review records for a 90-day sample period. Verify: (1) alerts were reviewed within the SLA the monitoring plan defines, (2) each alert has a corresponding triage record, (3) the review cadence matches the cadence the plan sets, (4) no alert was closed without an investigation record. (5) where agents select their own actions, the alert history includes departures from the authorised objective, and the triage record for each names the action taken on the run rather than closing on a note.
Post-market monitoring plan for a production AI system, held with its technical documentation, stating the field data intake, the analysis cadence and the conformity evaluation the data feeds.
Example: Post-market monitoring plan, classification-api v4, rev 3 dated 2026-07-04
Test: Select a production AI system and request its post-market monitoring plan. Verify: (1) the plan names the routes by which deployer and user performance data reaches the provider and the owner of each route, (2) field data from the last two review periods is present and analysed rather than only collected, (3) the analysis records a conclusion on whether the system still meets the requirements it was placed on the market against, with the evidence for that conclusion, (4) where the system runs alongside other AI systems, the interaction between them is analysed, (5) the plan's version matches the technical documentation version for the deployed release.
Questions (3)
Does each production AI system have a documented monitoring plan?
AI system behaviour degrades in ways not visible from infrastructure metrics alone. Monitoring must include AI-specific measures: confidence score distributions, null or refusal rates, output category distributions, in addition to standard latency and error rate metrics.
Which AI-specific metrics are included in your production monitoring for AI systems?
Programmes that monitor only latency and error rates are using generic application performance tooling, which misses the behavioural degradation patterns specific to AI systems. The last item applies only where agents select their own actions; where they do, it is the metric that shows an agent pursuing something other than the task it was given.
Which of the following does the monitoring plan record for each production AI system?
Options run from the most commonly in place to the least. Operational monitoring answers whether the system is healthy now; the field data limb answers whether what was claimed about it at release still holds. The conformity evaluation is the element most often missing, because collecting field data is easier than drawing a conclusion from it.