Performance Roadmap

Nine levels from measuring before you touch anything, through reading numbers honestly and finding where time goes, to the physics of loaded systems, the data layer, capacity, SLOs and incident method.

0 / 119 roadmap lessons mastered0%

Level 1 · Measure before you touch anything

The discipline that separates diagnosis from guessing: a symptom, a signal that confirms it, and a hypothesis written down before you look. Learn what each signal type can and cannot tell you, and the four numbers that describe any service.

Observability Is Not a Dashboard
Measure Before You Optimize
From Symptom to Root Cause
Metrics, Logs, Traces, Profiles
The Four Golden Signals
RED: Rate, Errors, Duration
USE: Utilization, Saturation, Errors
Instrumentation: From Code to Signal
OpenTelemetry Concepts

Level 2 · Read the numbers honestly

Averages hide outages and a user id in a label can take down the metrics backend. Counters, gauges and histograms as different questions, percentiles as the only honest summary of latency, and logs structured well enough to reconstruct a failure.

Four Metric Types, Four Questions
Counters: The Slope Is the Signal
Gauges: Blind Between Scrapes
The Average Was Fine and Users Were Not
Percentiles: Which One, and How Many Users Is That?
Cardinality: The Label That Took Down Monitoring
Label Sets That Survive a Year
Structured Logging: Fields a Program Can Read
Log Levels Are a Convention, Not a Standard
Correlation IDs: Turning Lines Into a Story
The Log Bill and What It Is Buying

Level 3 · Find where the time goes

Tracing locates time across services; profiling locates cost inside a process. Read a waterfall, find the critical path, spot N+1 by its shape, and know which of the two tools answers the question in front of you.

Where the Request Actually Went
Trace, Span, Attribute, Status
Parents, Children and Links
Carrying the Trace Across the Gap
Reading the Waterfall
The Critical Path Is the Only Path That Pays
The Comb: N+1 as a Visible Shape
Sampling Without Throwing Away the Evidence
When the Trace Runs Out of Answers
Self Time, Total Time, and Where the CPU Went
Reading a Flame Graph
Allocation Rate Is a Cost Even Without a Leak
Computing or Waiting?

Level 4 · Understand why loaded systems get slow

The physics. Queueing makes systems slow long before they fail, tail latency dominates user experience at scale, and Little's Law sizes every pool you will ever configure. This is the level that changes how you read every dashboard.

Latency Is a Distribution, Not a Number
Tail Latency: Why p50 Being Fine Does Not Help
Latency Budgets: Spending 200 Milliseconds on Purpose
Little's Law as Working Intuition
Queueing: Why Systems Get Slow Before They Get Broken
Saturation: The Reading Utilization Cannot Give You
Timeouts: The Latency Contract Nobody Writes Down

Level 5 · Identify the constrained resource

CPU, memory, disk and network, each with the signal that identifies it as the constraint — and the difference between a memory leak and a cache nobody bounded.

What "CPU Is At 60%" Actually Means
CPU Saturation: When Cores Become the Queue
Algorithmic Cost in a Request Handler
Memory Leaks: Growth That Does Not Come Back

Level 6 · Diagnose the data layer

Databases, caches and queues are where most production latency actually lives. Slow-query workflow, lock waits with idle CPU, pool saturation, hit rates that lie, stampedes, hot keys, backlogs and retry storms.

Which Signal Actually Means "The Database Is Slow"
The Slow Query Workflow
An Index Scan Is Not Automatically Faster
When the Join Strategy Is the Bottleneck
Low CPU, High Latency: Lock Contention
Replication Lag: Reads That Are Correct and Stale
A 95% Hit Rate Tells You Almost Nothing
Cache Stampede: Everyone Misses at Once
Hot Keys: When Aggregate Metrics Hide a Saturated Node
Six Queue Signals, Two That Wake You Up
The Backlog Arithmetic: Four Levers and a Drain Time
Depth Is Not an Emergency; Age Is
Retry Storms: The Load You Generated Yourself
Twenty Workers, All Busy, Five Hundred Waiting

Level 7 · Know your runtime and your users

Garbage collection, event loops, interpreters and JIT warm-up — labelled per runtime because none of it generalizes. Then the user's half of the budget: the browser waterfall, bundle cost and the vitals that measure experience.

Event-Loop Lag: One Callback, Everybody Waits
JavaScript Runtime Performance: V8 Where It Costs
CPython Performance: The Interpreter Tax and the GIL
C++ Memory Performance: Allocation, Copies and Locality
The Half of the Budget You Cannot See From the Server
Core Web Vitals as Signals, Not Scores
Reading the Browser Waterfall
JavaScript Costs Four Times, Not Once
Images: The Largest Bytes, Rarely the Largest Block
Layout, Paint and the Main Thread

Level 8 · Produce numbers that mean something

Load-test shapes that answer different questions, coordinated omission, benchmark hygiene, and telling a real regression from noise. Then capacity: headroom, autoscaling that arrives in time, and what a request actually costs.

Load Testing: What Question Is This Test Answering?
Load Test Shapes: The Shape Is the Hypothesis
Coordinated Omission: When the Load Generator Lies
Benchmarking: Does This Number Answer My Question?
Benchmark Fallacies: Confident Numbers That Are Wrong
Microbenchmark or End-to-End: Why p99 Did Not Move
Regression or Tuesday? Telling a Real Change from Noise
Capacity Planning: Traffic to Machines
Headroom: The Capacity You Deliberately Do Not Use
Autoscaling: Scaling on the Right Signal
Autoscaling Lag: The Gap Where the Outage Lives
Cost per Request: The Other Performance Metric
Capacity or Efficiency: Which Problem Are You Solving?

Level 9 · Run it in production, at scale

SLOs that describe user experience, alerts worth waking up for, error budgets as a decision tool — then distributed tail amplification, cross-region physics, incident method, and the trade-offs that make a system faster but worse.

SLIs: Measuring What the User Actually Feels
SLOs: A Target, a Window, and a Reason
SLAs: The Promise With Money Attached
Error Budgets: Unreliability You Are Allowed to Spend
Alerts Worth Waking Someone For
Alert Fatigue: The Page Nobody Reads
Burn-Rate Alerts: How Fast Is the Budget Going?
Dashboards Built Around Questions
What Changes When Work Crosses a Machine
Fan-Out: Waiting for the Slowest of Seven
Sequential or Parallel: Same Work, Different Latency
Cross-Region Latency Is Physics, Not Configuration
Packet Loss Buys You a Timeout, Not a Retransmit
Agreement Costs Round Trips
Debugging an Incident in Progress
Correlation Is Not the Root Cause
The Bottleneck Moves After Every Fix
Every Optimization Buys Something and Sells Something
Performance and Observability Anti-Patterns
Where an Agent Run Actually Spends Its Time
Inside One Model Call: Queue, First Token, Generation
Reading an Agent Run as a Trace
What One Agent Run Costs, and Which Term Dominates
The Agent Returned 200 OK and the Answer Was Wrong