Performance Interview Guide
Every question uses the same frame: what it tests, progressively stronger answers, green flags, red flags, follow-ups and a practical scenario. Strong candidates reason from measurement to mechanism — they do not name a fix in the first sentence.
An endpoint went from 100 ms to 2 seconds. How do you find out why?
CPU on your service is pegged at 100%. What do you investigate?
API p99 is 5 seconds. CPU is at 15%, memory is fine. What is going on?
Your cache hit rate is 90%. Is that good?
Why do teams look at p99 latency instead of the average? And when is p99 the wrong number?
Average latency is unchanged. p99 doubled overnight. What changed?
You have two weeks before a launch you expect to be 10× your current traffic. How do you load test it?
How would you define an SLO for a checkout service?
What deserves to wake a human at 3am, and what does not?
Memory grows about 50 MB per hour under stable traffic, and the process restarts every day or two. How do you investigate?
A queue is 2 million messages deep and growing. What do you do?
How many instances do we need for a traffic event four weeks from now?
When would you reach for a distributed trace instead of logs?
A teammate says their change makes a function 3× faster and wants to ship it. What do you ask?
You have one week to make this system meaningfully faster. How do you decide what to work on?
An AI agent takes 20 seconds per run and users are complaining. How do you make it faster?
Universal signals
Establishes a baseline, separates average from tail, finds the critical path, recognizes queueing, validates the change, leaves a regression guard.
Names a fix before a diagnosis, reads utilisation as saturation, trusts averages, trusts benchmarks, believes retries improve reliability.