AdvancedAutoscaling & Health← All practice

The Autoscaler Was Right and the Capacity Was Four Minutes Late

The report, in their words

Every weekday at 09:00 the service sheds requests for three to four minutes: latency climbs to six seconds, a few percent of requests time out, then everything recovers on its own. The autoscaling policy fires correctly every morning and adds the right number of instances. By the time they serve traffic the peak is over. The team has lowered the threshold twice and nothing changed.

Pull evidence

One item at a time, and nothing here tells you which one matters. Deciding what is worth looking at is most of the diagnosis.

One morning, hop by hop, with timestamps
Where the instance time goes
The shape of the morning ramp
Is CPU the right signal?
The scaling policy
What a cold instance does to latency
The cooldown
The database during the ramp
Where the fleet starts the day

0 of 9 inspected. You are not required to open all of them — a real investigation is judged on how few you needed.

What is your diagnosis?

Commit to one. Nothing below is shown until you do.

Guessing wrong and being told exactly why is the point of this page. Reading the answer first is not practice.