MON-010 Availability and SLO Monitoring with Status Communication
Description
Each customer-facing service has a service level objective for availability and for the latency of its principal operations, with the measurement window and the error budget that follows from it. The objective is measured continuously from signals that represent what a customer experiences. A breach or a projected breach raises an alert to a named responding team. A status channel readable without an account carries the current state of each service, is updated during a degradation within the time the incident process gives and records the start, the state changes and the resolution of each event. Each service also carries at least one qualitative target alongside its availability and latency objectives, written so that a miss is decidable. A missed target, quantitative or qualitative, produces a recorded corrective action inside the period the service level states. A revision to a published service level reaches the customers it binds before it takes effect. A development that puts a committed service level at risk, among them a change of control, a restructuring, the loss of a certification, the failure of a subcontractor and a decision to retire a service, is notified to the affected customers inside the notice period the agreement gives, through a named channel that is not the status channel, with the notice and its date recorded.
Rationale
Capacity monitoring tells the provider that a disk is filling. None of it tells a customer whether the service is working. An objective measured from the customer's side, with an error budget, turns availability from an aspiration into a number someone owns. A status channel stops the support queue from becoming the outage detector. INF-012 holds capacity and resource thresholds, MON-005 security alerting, INC-006 the breach notification and BCM-003 the recovery objectives that apply once an outage becomes a disruption. Issued in S4 from G-MON-3 in docs/review-canonical-quality.md section 7.3. The status channel carries what is happening now. A development that threatens a service level has not happened yet, so it never appears there, which is why that notice runs through a named channel with a record rather than through a page customers read once they have noticed a problem. A qualitative target is a target only where a miss is decidable: support responsiveness, change quality and data accuracy each become one the moment the threshold is written down and stay opinions until then. A revision notice is the counterpart, because a service level a customer is measured against can otherwise be lowered between two readings of the same page.
Applicability (9 profiles)
The service level objective and the status channel are commitments of a hosted service.
Condition: deployment_model in on-prem, hybrid, edge, on-device, embedded, air-gapped
Where the deployer runs the AI system on infrastructure it controls, the availability objective and the status channel for that service are its own. Under cloud-saas or cloud-single-tenant the row is satisfied by the provider: the deployer holds the provider's service level objective and status channel and reviews them under VND-006.
Condition: deployment_model in cloud-saas, cloud-single-tenant
The service level objective and the status channel are commitments of a hosted service. Conditional since 1.1: the profile lists every deployment model and this row is a commitment of a hosted service, so a provider that publishes weights or runs on infrastructure the customer controls has no hosted service to carry it (ADR-046 amendment, 2026-09-16).
The service level objective and the status channel are commitments of a hosted service.
Condition: deployment_model in on-prem, hybrid, edge, on-device, embedded, air-gapped
Where the deployer runs the AI system on infrastructure it controls, the availability objective and the status channel for that service are its own. Under cloud-saas or cloud-single-tenant the row is satisfied by the provider: the deployer holds the provider's service level objective and status channel and reviews them under VND-006.
The service level objective and the status channel are commitments of a hosted service.
EX-92 written in S8 wave B (migration 057) for three of its four clauses: at least one qualitative target per service alongside the availability and latency objectives, a recorded corrective action inside the period the service level states when a target is missed, notice to the customers a published service level binds before a revision takes effect, and notice of a development that puts a committed service level at risk through a named channel that is not the status page. Art. 30(2)(e), Art. 30(3)(a) and Art. 30(3)(b) hold full. Recorded limit (product owner, 2026-09-14): a multi-tenant provider exposes availability and integrity signals through its status channel, which EX-92 widens, and does not expose a per-customer confidentiality or authenticity indicator. The customer's ongoing monitoring of those two properties under RTS 2024/1773 Art. 9(1) is met by the provider's assurance evidence, the VND-011 shared responsibility matrix and the APP-015 audit reports, rather than by a live indicator. No MON control is authored for it and the Art. 9(1) mapping stays partial, with the measures that apply when a service level is missed now stated.
The service level objective and the status channel are commitments of a hosted service.
Art. 7 of Implementing Regulation (EU) 2024/2690 turns availability measurement into a reporting trigger: complete unavailability for more than 30 minutes, and limited availability for more than one hour affecting the lower of 5 % of the service users in the Union or one million of them. Art. 3(3) fixes the denominator as the contracting customers plus the natural and legal persons associated with business customers, which is not the same as monthly active users and has to be derivable. For an organisation that also holds the managed-service-provider role, Art. 10 measures the managed service rather than the cloud service, with the users of the customer's systems as the denominator, so availability of what is operated on a customer's behalf has to be measurable on its own.
Framework Mappings (9)
| DORA-Art.30.2.c | Availability, authenticity, integrity and confidentiality of data | informative |
| DORA-Art.30.2.e | Service level descriptions and their revisions | full |
| DORA-Art.30.3.a | Full service level descriptions with performance targets | full |
| DORA-Art.30.3.b | Notice periods and reporting obligations of the provider | full |
| DORA-RTS-2024/1773-Art.9.1 | Monitoring measures and key indicators in the contract | partial |
| 8.16 | Monitoring activities | informative |
| NIS2-CIR-Art.7 | Significant Incidents for Cloud Computing Service Providers | partial |
| A1.2 | Environmental Protections, Software, Data Back-Up Processes, and Recovery Infrastructure | partial |
| CC7.2 | Monitors System Components for Anomalous Behavior | informative |
Evidence (3)
Objective and alert configuration in the monitoring platform showing each service's availability and latency objective, the measurement window, the error budget and the alert routing.
Example: Service level objective configuration export of 2026-08-19 covering eight customer-facing services, each with its availability target, latency target, 30 day window, error budget and the on-call rota it pages
Test: Export the objective and alert configuration. Verify: (1) every customer-facing service in the service catalogue has an objective, (2) each objective names a measurement window and an error budget derived from it, (3) the signal behind each objective is taken from a request path a customer uses rather than from a host health check, (4) breach and burn-rate alerts route to a named team rather than to an unattended mailbox. (5) each service carries at least one qualitative target alongside its availability and latency objectives, with the threshold that decides a miss.
Availability report for the last period showing measured availability and latency against the objective for each service and the error budget consumed. It also records the corrective action taken on each missed target and any revision published to a service level during the period.
Example: Service availability report for July 2026, dated 2026-08-05, showing measured availability per service against target, error budget consumed and the two events that consumed most of it
Test: Read the availability report for the last complete period. Verify: (1) it reports measured availability and latency per service against the objective rather than an aggregate figure, (2) the measurement method matches the configured objective, (3) services that consumed their error budget carry a recorded action, (4) the events named in the report appear in the status channel history. (5) every missed target, quantitative or qualitative, carries a corrective action recorded inside the period the service level states, (6) a service level revised during the period reached the customers it binds before the revision took effect.
Status channel history for a recent degradation showing the entries published, their timestamps relative to detection and the resolution entry. The record also carries the notices given ahead of an event, where a development put a committed service level at risk, with the customers reached and the date.
Example: Status page incident history for the degradation of 2026-06-27, showing the first entry 14 minutes after detection, four state changes and the resolution entry with a follow-up note
Test: Take the most recent customer-affecting degradation. Verify: (1) a status entry was published within the time the incident process gives, measured from detection rather than from confirmation, (2) the entry names the affected service and the effect a customer would see, (3) state changes and the resolution were published, (4) the channel is readable without an account, (5) the event matches an entry in the availability report for the same period. (6) a development in the period that put a committed service level at risk carries a notice through a named channel other than the status channel, sent inside the notice period the agreement gives, with the customers reached and the date recorded, or the record shows that none arose.
Questions (3)
Does every customer-facing service have a defined availability objective?
An objective is a target with a measurement window, not a phrase in a sales contract. Answer no where availability is watched but no target is written down.
Which of the following are in place for availability monitoring?
Measurement from the customer's side means a probe or a request-level metric from the path a customer uses. Host uptime and container health checks answer a different question and a service can pass both while customers see errors. A qualitative target counts only where a threshold decides the miss. The last item is the one no status page carries: a change of control, a lost certification, a failing subcontractor or a retirement decision all threaten a service level before any event opens.
How are customers told about a degradation in progress?
Options run from strongest to weakest. The defined time is measured from detection, not from the point the cause is understood, because a customer needs to know the service is affected before anyone knows why.