Performance Roadmap

One track, nine levels. Start at Level 1 with measuring before you touch anything; every stage names what it needs first and what you should be able to do before moving on. Progress is stored locally in your browser.

Where to start

0 / 119 lessons masteredNot started 119Learning 0Practicing 0Mastered 0
  1. 1

    Level 1 · Measure before you touch anything

    Start here
    0/9

    The discipline that separates diagnosis from guessing: a symptom, a signal that confirms it, and a hypothesis written down before you look. Learn what each signal type can and cannot tell you, the four numbers that describe any service, and how instrumentation reaches a backend at all. Everything after this assumes you measure first.

    Before moving on: Turn a vague complaint into a symptom, the signal that would confirm it and a written hypothesis, and say whether a metric, a log or a trace answers the question in front of you.

  2. 2

    Level 2 · Read the numbers honestly

    0/13

    Averages hide outages and a user id in a label can take down the metrics backend. Counters, gauges and histograms as different questions, percentiles as the only honest summary of latency, and logs structured well enough to reconstruct a failure. It comes second because every later stage reads a p99 or a log line and needs you to read it right.

    Before moving on: Read a p50 / p99 pair and say which one the user feels, pick counter, gauge or histogram for a new measurement, and reject a label that would explode cardinality.

  3. 3

    Level 3 · Find where the time goes

    0/14

    Tracing locates time across services; profiling locates cost inside a process. Read a waterfall, find the critical path, spot N+1 by its shape, and know which of the two tools answers the question in front of you. These are the two tools every diagnosis from here on reaches for.

    Before moving on: Read a trace waterfall, mark the critical path, spot an N+1 by its shape, and decide from a flame graph whether a process is CPU-bound or waiting on I/O.

  4. 4

    Level 4 · Understand why loaded systems get slow

    0/9

    The physics. Queueing makes systems slow long before they fail, tail latency dominates user experience at scale, and Little's Law sizes every pool you will ever configure. It sits after percentiles because the queueing curve only makes sense once you read latency as a distribution, and before the resource and data-layer stages because saturation is what those diagnose.

    Before moving on: Size a pool with Little's Law, explain why latency climbs long before utilisation reaches 100%, and set a timeout from a latency budget instead of a guess.

  5. 5

    Level 5 · Identify the constrained resource

    0/8

    CPU, memory, disk and network, each with the signal that identifies it as the constraint — and the difference between a memory leak and a cache nobody bounded. Profiling tells you where cost goes; this stage tells you which resource has run out, so it follows profiling and saturation.

    Before moving on: Name the constrained resource from its signal alone, and tell a memory leak from an unbounded cache by plotting memory against traffic over a full daily cycle.

  6. 6

    Level 6 · Diagnose the data layer

    0/15

    Databases, caches and queues are where most production latency actually lives. Slow-query workflow, lock waits with idle CPU, pool saturation, hit rates that lie, stampedes, hot keys, backlogs and retry storms. Every one of these is queueing or a constrained resource seen from outside the process, which is why it comes after both.

    Before moving on: Run the slow-query workflow to a fix, explain lock waits on an idle CPU, size a connection pool, and read queue depth and oldest-message age as two different alarms.

  7. 7

    Level 7 · Know your runtime and your users

    0/12

    Garbage collection, event loops, interpreters and JIT warm-up — labelled per runtime because none of it generalizes. Then the user's half of the budget: the browser waterfall, bundle cost and the vitals that measure experience. Both halves read a profile or a waterfall, so profiling and the resource stage come first.

    Before moving on: Attribute a tail spike to garbage collection or event-loop lag from the right runtime signal, and read a browser waterfall to say whether the page is slow or the API is.

  8. 8

    Level 8 · Produce numbers that mean something

    0/13

    Load-test shapes that answer different questions, coordinated omission, benchmark hygiene, and telling a real regression from noise. Then capacity: headroom, autoscaling that arrives in time, and what a request actually costs. A load test is queueing theory run on purpose and a capacity estimate is arithmetic on percentiles, so this waits for both.

    Before moving on: Pick the load-test shape that answers the question you have, run a benchmark that avoids the common fallacies, and state a capacity estimate with its headroom and its assumptions labelled.

  9. 9

    Level 9 · Run it in production, at scale

    0/26

    SLOs that describe user experience, alerts worth waking up for, error budgets as a decision tool — then distributed tail amplification, cross-region physics, incident method, agent latency budgets, and the trade-offs that make a system faster but worse. It comes last because an SLI is a percentile with a target, an incident is every earlier stage under pressure, and a trade-off is only visible once you can measure both sides.

    Before moving on: Write an SLI and SLO for one user journey, set a burn-rate alert, and run an incident from the pivotal signal to a review that names the regression test and the new alert.