INF-012 Capacity and Performance Management
Description
Current and projected resource utilisation across compute, storage, network and API throughput is monitored against defined thresholds. Capacity plans are maintained so that systems can meet operational demand. Alerts are configured for resource exhaustion conditions before they affect service availability. Log storage is provisioned to hold logs for their required retention period and its utilisation is monitored against the same thresholds.
Rationale
Resource exhaustion, whether from organic growth or denial-of-service, is a primary availability threat. Proactive capacity management enables timely scaling decisions and preserves SLA commitments.
Applicability (9 profiles)
Framework Mappings (12)
| I&S-02 | Capacity and Resource Planning | full |
| I&S-02 | Capacity and Resource Planning | full |
| 8.6 | Capacity management | full |
| AML.M0036 | Limit AI Workload Resource Consumption | informative |
| NIS2-CIR-13.1 | Supporting Utilities | informative |
| NIS2-CIR-4.2 | Backup and Redundancy Management | informative |
| AU-4 | Audit Log Storage Capacity | full |
| SC-5 | Denial-of-service Protection | informative |
| SC-6 | Resource Availability | full |
| MS-2.12-003 | AI Environmental Impact Assessment | MS-2.12-003 | informative |
| LLM06 | Unbounded Consumption | informative |
| A1.1 | Capacity Management | full |
Evidence (3)
Resource utilisation dashboards and alerting configuration for production infrastructure showing monitoring coverage and defined threshold alerts for compute, storage, network, and API throughput.
Example: AWS CloudWatch, Datadog, or Grafana dashboard export showing CPU, memory, storage, and API latency metrics for production workloads, with alert threshold configuration visible
Test: Request a monitoring dashboard export and the alert configuration for capacity thresholds. Verify: (1) metrics are collected for all critical resource types (CPU, memory, disk, network I/O, API throughput); (2) alert thresholds are set below resource exhaustion limits; (3) alerts are routed to an active on-call channel or queue; (4) review the alert history, confirming that alerts fired before actual exhaustion events in the last 90 days.
Capacity plan or capacity review record demonstrating projected utilisation has been assessed against operational demand forecasts.
Example: Quarterly capacity review report or capacity planning record (e.g., Confluence page or document), showing projected versus actual utilisation trends and planned scaling actions
Test: Request the most recent capacity plan or review record. Verify: (1) the plan covers all critical infrastructure components; (2) current utilisation is compared against projected demand; (3) scaling decisions or actions are documented; (4) the review was completed within the defined frequency.
Log storage capacity metric history showing storage utilisation trends and any pipeline failure events and their resolution.
Example: AWS CloudWatch or equivalent metrics export showing log bucket storage utilisation and log delivery failure rate over the past 90 days
Test: Query storage capacity and pipeline health metrics for the last 90 days. Verify: (1) storage utilisation has not reached or exceeded the defined alert threshold without triggering an alert; (2) any pipeline failure events have a corresponding incident or remediation record; (3) no log loss events have occurred without detection.
Questions (3)
Is current and projected resource utilisation monitored against defined thresholds?
Monitoring should cover all critical resource types. Alerts should fire with enough lead time to allow scaling decisions before service is impacted.
How frequently is capacity planning reviewed to ensure production infrastructure can meet operational demands?
Auto-scaling removes much of the risk but does not eliminate the need for capacity planning at the service limit level. Quarterly reviews are a minimum for services with defined availability SLAs.
Which of the following are monitored against a defined threshold with an alert configured before exhaustion?
Options run from the most commonly monitored to the least. Log storage is the one that fails quietly: the store fills, the oldest events roll off and the retention the organisation believes it holds is gone before anyone looks.