DevOps Roadmap
Nine stages in one order, starting at Source to artifact. Every stage names what it needs first and what you should be able to safely do before moving on — the order is load-bearing, because a practice built on an unstable artifact cannot be made safe further down the pipeline. Progress is stored locally in your browser.
Where to start
DevOps / Production Engineering
9 stages · 0/185 lessonsTurning source into a running production system, operating it safely, changing it continuously, and recovering when it breaks.
- Source to artifact
- A runnable unit and the things that vary around it
- Getting change into production safely
- Infrastructure you can reproduce
- Orchestration, discovery and traffic
- Operating it, and responding when it breaks
- Capacity, cost and the stateful parts
- Release engineering, supply chain and platform
- Multi-region, recovery and production engineering
- 10/18
Source to artifact
Start hereWhere the loop begins: what makes production different, why a commit is a candidate for production state rather than history, and how CI, builds and registries turn that commit into one immutable, addressable artifact. It comes first because everything after it assumes the artifact is the unit that moves.
Before moving on: Answer "what exactly would I deploy, and which commit produced it" with a digest, and explain why rebuilding per environment destroys that evidence.
- What Production Engineering Is
- The Production Loop
- Source Control as Production Infrastructure
- Trunk-Based Development
- Protected Branches
- Required Checks
- Continuous Integration
- CI Is a Feedback System
- Designing the Pipeline
- Flaky Tests
- What a Build System Actually Is
- Reproducible Builds
- Dependency Pinning
- What an Artifact Is
- Build Once, Deploy Many
- Tags Versus Digests
- Semantic Versioning, and Where It Stops Applying
- Artifact Registries
- 20/19
A runnable unit and the things that vary around it
Packaging the artifact so it runs the same way everywhere, then supplying the parts that legitimately differ per environment — configuration and credentials — without rebuilding it. Container layers and the process and signal model, config validated at startup, workload identity instead of pasted secrets, and why staging is not production.
Before moving on: Ship artifact plus config instead of "the code": build a small image, say what happens to it on
SIGTERM, and place each value in the image, the config or the secret store.Needs first:Source to artifact- The Container Lifecycle
- Image Versus Container
- Layers and the Build Cache
- Multi-Stage Builds
- What Image Size Actually Costs
- PID 1 and Signals
- Graceful Shutdown
- Artifact Plus Configuration
- Build-Time and Runtime Configuration
- Validate at Startup, Fail Clearly
- A Config Change Is a Production Change
- What Counts as a Secret, and Where It Must Not Be
- Secret Managers and What They Actually Give You
- Workload Identity
- Secrets in CI
- What an Environment Is For
- Parity That Is Worth Paying For
- Environment Drift
- Promotion Between Environments
- 30/18
Getting change into production safely
Deployment stops being an event and becomes a routine. Release versus deployment, the pipeline that carries one artifact through environments, the strategies — recreate, rolling, blue/green, canary, flags — with what each costs and how you get back, and blast radius as the idea that organises all of it. It needs a stable artifact and config to move, which is why it sits after the first two stages.
Before moving on: Put a new version in front of real traffic in a shape you chose, explain why version coexistence is the hard part, and decide whether rolling back or rolling forward is safe for a given change.
- Deployment Is Not Release
- Continuous Delivery
- Continuous Deployment
- The Deployment Pipeline
- Deployment Strategies
- Recreate: Stop Everything, Then Start the New Thing
- Rolling: Two Versions, One Database
- Blue/Green: Paying for the Fastest Rollback There Is
- Canary: One Percent, Then Five, Then Watch
- Feature Flags: Deploy Is Not Release
- Progressive Delivery: Exposure as a Dial
- Version Coexistence: N and N+1, in Both Directions
- Canary Analysis: Compared Against What?
- A Successful Deploy Is Not Evidence of a Healthy System
- Rollback: Only Useful If It Is Actually Safe
- Roll Forward: When Going Back Is the Harder Option
- Change Size: Why Small Changes Are Safer, and When They Are Not
- Blast Radius: If This Is Wrong, How Much Does It Affect?
- 40/13
Infrastructure you can reproduce
The environment the artifact runs in becomes code that can be reviewed, planned and diffed rather than a console someone clicked. Plans, state, drift, the destructive change a rename hides, immutable infrastructure and preview environments — and, because the console is still there, who is allowed to touch production by hand and with what privilege.
Before moving on: Stand an environment up from a repository, read a plan and say what it will destroy before it does, and detect when reality has stopped matching the code.
- Infrastructure as Code
- Declarative vs Imperative Infrastructure
- The Plan: Desired vs Current
- State
- Drift
- Destructive Changes: What a Rename Really Does
- Immutable Infrastructure
- Pets and Cattle, Read Carefully
- Preview Environments
- Ephemeral Environments
- Manual Production Changes
- Production Access
- Least Privilege in Production
- 50/20
Orchestration, discovery and traffic
Running many replicas of many services, letting them find each other, and putting traffic in front of them without dropping requests. Kubernetes is taught as one implementation of the scheduling and reconciliation problem — pods, deployments, requests and limits, probes — and then the operational half of the network: DNS under change, load balancer health, connection draining and certificate lifecycles.
Before moving on: Explain reconciliation and the scheduler well enough to diagnose an OOM kill, a failing probe or a pending pod, and drain a replica out of a load balancer without losing a request.
- Do You Need Kubernetes?
- The Problems Kubernetes Answers
- Cluster, Control Plane, Nodes, Pods
- Pods: The Unit That Gets Scheduled
- Deployments: Declaring What Should Be Running
- ReplicaSets: The Layer You Should Not Manage
- Services: A Stable Address Over Moving Pods
- Getting Traffic Into the Cluster
- Reconciliation: The Loop Under Everything
- Apply Is Not Running
- The Scheduler, and Why a Pod Is Pending
- Requests and Limits
- Probes: Readiness, Liveness and Startup
- ConfigMaps and Secrets
- Service Discovery in Operation
- DNS in Production
- Operating a Load Balancer
- Draining: Stopping Without Dropping
- Certificates as an Operational Object
- Renewal: Automating the Thing That Expires
- 60/20
Operating it, and responding when it breaks
What an operator does with observability: alerts that page only for symptoms a human must act on, dashboards annotated with deploys, on-call that stays healthy, and the incident loop from detection through mitigation to a postmortem whose action items change the system. It sits here because you need something deployed and running before there is anything to operate — and because "what changed?" is only answerable once deploys leave a trail.
Before moving on: Tell whether production is healthy without asking anybody, run an incident from detection to mitigation, and write a postmortem whose action items change the system rather than the people.
- Using Observability, Not Building It
- An Alert Should Demand Action
- Alert on Symptoms, Not on Causes
- Alert Fatigue
- Dashboards an Operator Can Act On
- Deploys on the Same Timeline as the Symptom
- On-Call Is Production Ownership
- Rotations People Can Sustain
- What Happens Between the Page and the Postmortem
- Severity: What It Should Reflect
- Stop the Harm Before You Understand It
- Roles During an Incident
- Reconstructing What Actually Happened
- Telling People What Is Happening
- Postmortems
- Root Cause vs Contributing Factors
- Action Items That Change the System
- Production Debugging
- Deployment-Centric Debugging
- Change Correlation
- 70/27
Capacity, cost and the stateful parts
How much traffic the system can take, what happens when it takes more, what each unit of work costs, and how to change a schema without an outage. Headroom and load shedding, autoscaling signals and their failure modes, cost drivers, and expand/migrate/contract with the backfills and locks that make a migration the change most likely to cause an outage. Reliability, money and the database stop being separate conversations.
Before moving on: Build a capacity model with headroom for failure and deploys, pick an autoscaling signal that reflects the real constraint, and run a zero-downtime migration as one event coupled to its deploy.
- Capacity Management
- Building a Capacity Model
- Headroom
- Load Shedding
- Overprovisioning
- Idle Capacity
- Autoscaling
- Choosing the Scaling Signal
- How Autoscaling Fails
- Horizontal Pod Autoscaling
- Queue-Based Autoscaling
- Scale to Zero
- Cost Awareness
- Cost Drivers
- Cost Per Request
- FinOps
- Operating a Production Database
- The Connection Budget
- Why Migrations Are the Dangerous Change
- Expand, Migrate, Contract
- Zero-Downtime Migrations
- A Migration and a Deploy Are One Event
- Backfills
- Destructive Migrations
- Why Stateful Workloads Are Harder
- Operating Queues and Scheduled Work
- Operating a Cache
- 80/25
Release engineering, supply chain and platform
Delivery becomes a product with users. Release manifests, change management and an audit trail that answers "what is in production, where did it come from and who approved it"; provenance, signing, SBOMs and scanning for everything between a dependency and a running artifact; and golden paths, guardrails and policy as code that make the safe path the easy one. It builds on the artifact, the pipeline and the infrastructure code of the earlier stages.
Before moving on: Answer what is in production with evidence rather than recollection, verify an artifact's provenance before promoting it, and design a golden path other teams choose over doing it by hand.
- Release Engineering as a Discipline
- The Release Manifest
- Change Management
- The Audit Trail
- Promotion
- Artifact Retention
- The Delivery Chain as Attack Surface
- Build Provenance
- The Builder Is Inside the Trust Boundary
- Signing and Verifying Artifacts
- Software Bill of Materials
- Scanning, and Why a Finding Is Not a Risk
- Securing the Pipeline Itself
- CI Security
- Platform Engineering
- The Internal Developer Platform
- Golden Paths
- Service Templates
- Self-Service Infrastructure
- Guardrails, Not Gates
- Policy as Code
- Developer Experience as an Operational Metric
- How to Automate Something
- The Automation Trap
- Toil
- 90/25
Multi-region, recovery and production engineering
Losing a region, a database or a dependency and still having a tested procedure back to service. Backups you have actually restored, RTO and RPO tied to real runbooks, region failover and the capacity question it raises, readiness reviews and service ownership, break-glass access, and the same discipline applied to agent systems, where prompts and models are deployable inputs. Last because it engineers the properties that let an organisation keep many services up, not one.
Before moving on: Run a restore drill and a region failover inside objectives you have measured, review a service for readiness before it carries traffic, and roll out a prompt or model change behind a canary with a kill switch.
- Backup Operations
- Restore Drills
- Disaster Recovery as an Operation
- RTO and RPO
- Region Failover
- Operating in More Than One Region
- Capacity During Failover
- Partial and Logical Data Recovery
- Production Readiness Review
- The Readiness Scorecard
- The Ownership Record
- Runbooks
- Runbook Anti-Patterns
- Break-Glass Access
- Access Review
- How Networks Fail in Production
- Reducing Blast Radius
- Production Anti-Patterns
- CI/CD Anti-Patterns
- Learning Across Incidents
- Deploying an Agent
- Prompts and Models Are Deployables
- Canarying a Model or Prompt Change
- The Agent Kill Switch
- Agent Cost in Production