Skip to content

Type to search pages.

View .md

Runbooks

Eleven alert rules watch the hosted control plane. Each one names its runbook in its own description, so the page arrives in the email and the SMS rather than having to be found. This is the index of those pages.

They are created by infra/terraform/modules/alerting and they are off by default. Production turns them on. Staging does not, and that is deliberate: staging is where a bad deploy is supposed to be caught, so it breaks on purpose several times a week. A page for that is a page somebody learns to ignore, and it is the same page production sends.

Alert Severity Runbook
unreachable 0 The control plane is unreachable
database-unreachable 0 The database is not answering
server-errors 1 Server errors
restart-loop 1 Revision health
bootstrap-job-failed 1 A job failed
maintenance-job-failed 1 A job failed
replicas-below-minimum 2 Revision health
database-storage 2 Database storage
database-connections 2 Database connections
database-cpu 3 Database CPU
certificate-expiring 3 The certificate

Each name is prefixed with the stack’s own, so the production rule for the first row is afcpprod-unreachable.

Severity 0 means the service is down for customers. Severity 1 means it is failing and probably visible. Severity 2 and 3 are warnings with hours or days in them, and neither should be looked at before the sun is up.

One more control lives outside Azure and pages through GitHub instead: the vulnerability scan.

One action group, with an email receiver and an optional SMS receiver. The addresses are not in this repository. They are passed as TF_VAR_alert_emails, TF_VAR_alert_sms_country_code and TF_VAR_alert_sms_number, because a plan runs on every pull request into a step summary that is world readable, and an address in a variable file leaves through a diff.

Enabling alerting with no receiver fails at plan. An action group with no receivers creates cleanly, attaches to every rule, reports healthy, and delivers nothing to anybody. That is worse than no alerting, because it looks like alerting.

The engine. Nothing here watches a customer’s own continuous integration. The engine runs in their infrastructure and reports through ingestion, and its own alert rules are in observability/alerts/antifailure.rules.yml for anybody running Prometheus.

The application’s own counters. GET /metrics exposes what the process counted itself, and Azure Monitor cannot read it. The operations page is the guide to those, and it is the page to open second on any incident that starts here.

Anything outside Azure. The availability test runs from Microsoft managed agents in other regions. That is outside this stack, its group, its region and its network, and it is not outside Azure. A failure large enough to take the prober and the service together reports nothing at all.