Backend Engineering Roadmap
Start at Request in, response out. Nine stages in one order; every stage names what it needs first and what you should be able to build before moving on, and each assumes the failure modes of the ones before it. Progress is stored locally in your browser.
Where to start
Backend engineering
9 stages · 0/200 lessonsWhat happens from the moment a request reaches the service until data is returned, persisted, queued, cached, retried, observed and served at scale.
- Request in, response out
- Structure, data access and who is calling
- Transactions, queries, caches and other people's APIs
- Work that outlives the request
- Timeouts, retries, idempotency and limits
- Observability, testing, security and performance work
- Scaling out and shipping without dropping requests
- Distributed patterns, failure handling and multi-tenancy
- Production backend engineering
- 10/17
Request in, response out
Start hereThe request lifecycle with the framework removed: what a backend is responsible for because the client cannot be trusted with it, how a socket becomes a request object, how a method and a path pick a handler, the middleware pipeline every request passes through, and where malformed input is rejected. Everything later assumes you can see this path end to end.
Before moving on: Build an HTTP service that rejects malformed input at the edge, runs a handler and returns the right status code, and explain what happened between the socket and the response body without naming a framework.
- What a Backend Actually Is
- The Request Lifecycle
- What the Backend Is Responsible For
- The Trust Boundary
- Anatomy of an HTTP Server
- Request and Response Objects
- Status Codes From the Server's Side
- Request Bodies and Streaming
- How a Route Becomes a Function Call
- Path Parameters
- Query Parameters
- Route Precedence
- What a Handler Is Responsible For
- The Middleware Pipeline
- The Three Validations
- Transport Validation
- Reporting Validation Failures
- 20/19
Structure, data access and who is calling
Once a handler works, the question is where the rest of the code goes. Service and repository layers keep handlers thin and business rules testable, business validation is separated from the transport validation of the first stage, and the data access layer — ORM, query builder or raw SQL — is chosen per query, with a look at what an ORM actually does under the method call. Then authentication establishes who is calling and authorization decides what they may touch, down to the object-level check whose absence is the most common serious backend vulnerability. This is where a personal project becomes something other people can safely log into.
Before moving on: Build a service with thin handlers, business rules in a testable layer and an object-level authorization check on every resource, and explain why authentication and authorization are different questions.
Needs first:Request in, response out- Transport, Application, Domain, Infrastructure
- The Service Layer
- The Repository Layer
- Fat Controllers
- Dependency Management Without the Container
- Business Validation
- Parse, Do Not Validate
- Choosing a Data Access Layer
- What an ORM Actually Does
- Middleware Ordering Is a Correctness Decision
- Authentication in a Backend
- Credentials and Password Handling
- Session Authentication
- Token Authentication and the Revocation Problem
- Authorization in Backends
- Authentication vs Authorization
- Where the Check Belongs
- Role-Based Access Control
- Object-Level Authorization
- 30/20
Transactions, queries, caches and other people's APIs
The data path under load. Transaction boundaries and why a network call inside one is a resource problem; ORM, query builder and raw SQL as an engineering decision; the N+1 pattern and the connection pool that is the real concurrency limit; cache-aside with invalidation and TTLs, and the cases where a cache buys nothing; a timeout on every call that leaves the process; and DTOs so the database row is not what the client sees. It builds on the repository layer from the previous stage.
Before moving on: Decide what belongs in one atomic unit, write list endpoints whose query count does not grow with the result set, add a cache you can invalidate, and put a timeout on every external call.
Needs first:Structure, data access and who is calling- Transactions from Application Code
- Where the Transaction Boundary Goes
- One Transaction or Two
- External Calls Inside a Transaction
- What an ORM Buys and What It Costs
- The N+1 Query Problem
- Eager Loading and Batching
- Query Builders
- Raw SQL in Application Code
- Connection Pools
- Caching in Backends
- Cache-Aside
- Cache Invalidation
- TTL and Expiry
- When Not to Cache
- Calling Something You Do Not Control
- Timeouts
- Serialization: Objects to Bytes
- Three Models, Not One
- Schema Leakage
- 40/19
Work that outlives the request
Work that does not belong in the request path. Deciding what to defer, the queue lifecycle from enqueue to dead-letter, commands versus events and the consumers that act on them, inbound webhooks verified on the raw payload, and uploads that go straight to object storage through presigned URLs. It comes after the data stage because a job is a transaction that finishes later, and the recurring question is what the user sees while the work is still running.
Before moving on: Move a slow step off the request path onto a queue with a dead-letter policy, verify an inbound webhook signature on the raw body, and accept a file upload that never passes through your process.
- Request or Background?
- Background Jobs
- Job Queues
- Queue Semantics
- Scheduled Jobs
- Dead-Letter Queues
- Commands vs Events
- Event-Driven Backends
- Naming Events
- Writing Event Consumers
- Inbound Webhooks
- Webhook Signature Verification
- Outbound Webhooks
- Choosing an Upload Path
- File Uploads Through the Backend
- Presigned URLs
- Object Storage
- What Happens After the Bytes Land
- Serving Files
- 50/20
Timeouts, retries, idempotency and limits
The stage that makes the previous two safe. Retries with backoff and jitter, circuit breakers and bulkheads for a flaky dependency, and the rate limits you enforce; idempotency keys so a client that presses the button twice is charged once; at-least-once delivery and duplicate detection for jobs and webhooks; and two requests on one row, resolved with optimistic versioning, pessimistic locks or atomic operations. Every retryable path becomes a safe-to-retry path, and every unbounded thing gets a bound.
Before moving on: Make every retryable path safe to retry with an idempotency key, bound a flaky dependency with a timeout, backoff and a circuit breaker, and choose optimistic or pessimistic locking for a two-requests-one-row race.
- Retries
- Backoff and Jitter
- Circuit Breakers
- Bulkheads
- Rate Limiting
- Rate Limit Algorithms
- Idempotency in Backends
- Idempotency Keys
- The Idempotency Key Flow
- Idempotency Storage
- At-Least-Once Delivery
- Duplicate Detection
- Job Idempotency
- Webhook Idempotency
- Webhook Retries and Ordering
- Backend Races
- Optimistic Concurrency
- Pessimistic Locking
- Atomic Operations
- Resource Limits
- 60/22
Observability, testing, security and performance work
Operating the service you can now build. An error taxonomy and error boundaries that stop internals leaking; correlation ids, structured logs, metrics, traces and health checks so you can answer questions you did not anticipate when you deployed; configuration and secrets kept out of the code; a test strategy chosen by what each layer can prove, including integration tests against a real database and performance tests that measure more than one requests-per-second number; and the security checklist every service owes, from SQL injection to SSRF.
Before moving on: Follow one slow request through logs, metrics and traces, prove a change is safe with an integration test against a real database, and pass a security review of the service itself rather than of the network in front of it.
Needs first:Timeouts, retries, idempotency and limits- An Error Taxonomy That Maps Cause to Response
- Error Boundaries: Three Translations, Not One
- Not Leaking Your Internals
- Correlation Ids That Survive Every Hop
- What a Backend Should Actually Log
- Structured Logging
- The Metrics a Backend Must Emit
- Tracing From the Backend's Side
- Health Checks: Startup, Readiness, Liveness
- Configuration: Separating Code From Environment
- Secrets Are Not Configuration
- Validate at Startup, Fail Loudly
- A Test Strategy Chosen by What Each Layer Can Prove
- Test Against the Real Database
- Contract Tests Between Services
- Performance Testing a Backend
- The Backend Security Checklist
- SQL Injection
- Command Injection
- SSRF — When the Backend Fetches a URL
- Secrets in Logs
- What Serialization Costs
- 70/20
Scaling out and shipping without dropping requests
More than one copy of the service. Statelessness is what makes horizontal scaling and load balancing possible; then read replicas, pagination and autoscaling signals, and shipping without dropping requests: containers, graceful shutdown, rolling deploys and expand-contract migrations that two versions read at once. The runtime-model lessons sit here because the runtime decides what one instance can handle before you add a second.
Before moving on: Run the service behind a load balancer with no in-memory state, deploy a new version while the old one is still serving, and migrate a schema that both versions read at the same time.
- Stateless Services
- Making an Existing Service Stateless
- Horizontal vs Vertical Scaling
- Load Balancing, From the Backend's Side
- Sticky Sessions
- Autoscaling a Backend
- Read Replicas From the Application
- Pagination That Survives a Large Table
- Deployment Models
- Containerizing a Backend
- Graceful Shutdown
- Rolling Deployments
- Expand and Contract Migrations
- Schema Migrations from the Application Side
- Running Two API Versions in One Service
- Mapping Services Across Cloud Providers
- Running a Backend on Kubernetes
- Serverless Backends
- Choosing a Runtime
- Backend Runtime Models
- 80/21
Distributed patterns, failure handling and multi-tenancy
What changes when a function call becomes a network call. Monolith, modular monolith and microservices compared honestly; failure propagation and cascading failure; the dual-write problem and the transactional outbox; eventual consistency; multi-tenancy and tenant isolation; and the concurrency of a shared system: backpressure, cache stampedes, request coalescing and deadlocks. It needs the retry and idempotency stage as much as the scaling one, because every pattern here is a retry across a network.
Before moving on: Explain what happens when half of a write succeeds and fix it with a transactional outbox, stop one slow dependency from taking down services that do not depend on it, and isolate tenants so one deployment never leaks data between customers.
Needs first:Scaling out and shipping without dropping requestsTimeouts, retries, idempotency and limits- The Monolith
- The Modular Monolith
- Microservices
- Comparing Backend Architectures
- Synchronous vs Asynchronous Communication
- Failure Propagation
- Cascading Failure
- The Dual Write Problem
- The Transactional Outbox
- Eventual Consistency in Practice
- Keeping a Search Index in Sync
- Multi-Tenancy
- Tenant Isolation
- Attribute-Based Access Control
- Backpressure
- Worker Scaling
- Local vs Distributed Cache
- Cache Stampede
- Unbounded Concurrency
- Request Coalescing
- Deadlocks in Application Code
- 90/42
Production backend engineering
The rest of the domain, read with the judgement the earlier stages built: production debugging from symptom to cause, the runtime internals of Node, Python and worker processes, canary and blue-green deploys with feature flags, the low-level HTTP mechanics of accepting connections and parsing requests, the remaining structure and auth lessons, and agent-enabled backends, where a model choosing a tool is a client choosing an endpoint.
Before moving on: Take an API that was 100 ms yesterday and is 3 s today from symptom to cause, ship the mitigation behind a canary or a feature flag, and say what to change and what to leave alone.
Needs first:Distributed patterns, failure handling and multi-tenancyObservability, testing, security and performance work- Debugging a Backend in Production
- Why Is My API Slow?
- The Common Backend Failures
- Backend Code Smells
- Deploys Are the First Suspect
- Memory Leaks in Backend Services
- Retry Storms
- Connection Pool Exhaustion
- Queue Backlog
- Blocking the Event Loop
- The Node Event Loop
- Python Runtime Models
- Worker Processes
- C++ Backend Services
- Canary Deployments
- Blue-Green Deployments
- Feature Flags: Rollout, Kill Switches and Debt
- Defence in Depth
- Dependency Security
- Keep-Alive and Connection Reuse
- Accepting Connections
- Parsing HTTP
- Deserialization: Bytes to Objects
- Every Input Surface
- Database Constraints
- When the Repository Is Just Indirection
- Alternatives to Layering
- What Belongs in the Pipeline
- Request Context Propagation
- The Error Boundary
- Authenticate First, or Rate-Limit First?
- Where Sessions Live
- OAuth and OIDC From the Backend Side
- API Keys
- Email and Notifications
- Backend and Its Neighbours
- The Backend Reasoning Loop
- What an Agent Adds to a Backend
- A Tool Call Is a Backend Call
- Agent Authorization
- Budgets, Deadlines and Step Limits
- Agent Audit Logs