Alerting & Operational Signals
Using observability rather than building it: alerts that demand action, symptom-based paging, dashboards an operator can act on, and the cost of noise.
The operator's path from a page to a hypothesis: alert, dashboard, trace, logs — and what each hop is actually for.
The rule that decides what is allowed to page a human: if there is no action a person would take right now, it is not an alert.
"Users cannot check out" beats "CPU is 81%" — with the honest exception of infrastructure conditions that have a specific, immediate response.
Noise trains responders to ignore the pager, and the cost is paid on the one night the alert was real.
Built around the questions an incident asks, in the order it asks them — not around everything the system can emit.
The highest-signal overlay there is: a vertical line at each release, drawn across the error rate graph.