advanced
Autoscaling worked perfectly and the site still went down
A marketing email drives 6× traffic in 90 seconds; new instances arrive four minutes later, to a fleet that has already collapsed.
The page
CRITICAL · storefront-api · availability 61% (SLO 99.9%) · p99 timeout · 5xx 4,200/s
Timeline — in the order it was observed
Observation order is not causal order. The first thing anyone noticed is rarely the first thing that happened.
- 11:00:00Marketing sends a campaign email to 1.4 million subscribers. Engineering was not told.campaign log
- 11:00:40Request rate begins climbing from 1,200 req/s.lb request rate
- 11:02:10Request rate reaches 7,400 req/s — just over 6× baseline in 90 seconds.lb request rate
- 11:02:30Fleet CPU crosses 70%. The autoscaler is evaluating on a 60-second metric window.cpu_utilization
- 11:03:15p99 latency crosses 8 s. Requests begin timing out at the load balancer.request_duration
- 11:03:40Autoscaler issues a scale-out from 8 to 24 instances.autoscaler events
- 11:04:20Availability drops below 70%. The existing fleet is saturated and shedding requests.availability
- 11:07:50First new instances pass health checks and start receiving traffic — and are slower than the instances they were sent to help.per-instance latency
- 11:12:30Fleet reaches 40 instances and stabilises. Peak instance count never hit the configured maximum of 60.instance count
The system
Pull up evidence · 0/7 opened
Most of this evidence is consistent with several explanations. Keep going until something narrows it.
Autoscaling Lag: The Gap Where the Outage LivesAutoscaling: Scaling on the Right SignalHeadroom: The Capacity You Deliberately Do Not UseCapacity Planning: Traffic to MachinesJIT and Warm-Up: The First Thousand Requests Are a Different ProgramSaturation: The Reading Utilization Cannot Give YouQueueing: Why Systems Get Slow Before They Get BrokenLoad Test Shapes: The Shape Is the HypothesisConcurrency Limits: An Unbounded Server Is a Slower Server