DevOps Roadmap

Nine levels, each defined by what you can safely do once you have it rather than by what you have read. The order is load-bearing: every level assumes the failure modes of the one before it, and a practice built on an unstable artifact cannot be made safe further down the pipeline.

0 / 185 mastered0%
Level 1

Source to artifact

You can turn a commit into one trustworthy, immutable, addressable build output. Everything after this level assumes the artifact is the unit that moves; if the artifact is not stable, no later practice can be. At the end of this level you can answer "what exactly would I deploy, and which commit produced it".

What Production Engineering Is
The Production Loop
Source Control as Production Infrastructure
Trunk-Based Development
Protected Branches
Required Checks
Continuous Integration
CI Is a Feedback System
Designing the Pipeline
Flaky Tests
What a Build System Actually Is
Reproducible Builds
Dependency Pinning
What an Artifact Is
Build Once, Deploy Many
Tags Versus Digests
Semantic Versioning, and Where It Stops Applying
Artifact Registries
Level 2

A runnable unit and the things that vary around it

You can package that artifact so it runs the same way anywhere, and supply the parts that legitimately differ per environment — configuration and credentials — without rebuilding it. You stop shipping "the code" and start shipping artifact plus config, which is what actually determines behaviour.

The Container Lifecycle
Image Versus Container
Layers and the Build Cache
Multi-Stage Builds
What Image Size Actually Costs
PID 1 and Signals
Graceful Shutdown
Artifact Plus Configuration
Build-Time and Runtime Configuration
Validate at Startup, Fail Clearly
A Config Change Is a Production Change
What Counts as a Secret, and Where It Must Not Be
Secret Managers and What They Actually Give You
Workload Identity
Secrets in CI
What an Environment Is For
Parity That Is Worth Paying For
Environment Drift
Promotion Between Environments
Level 3

Getting change into production safely

You can put a new version in front of real traffic on purpose, in a shape you chose, and take it back. This is the level where deployment stops being an event and becomes a routine, and where you learn that the hard part is never the deploy — it is version coexistence and the way back.

Deployment Is Not Release
Continuous Delivery
Continuous Deployment
The Deployment Pipeline
Deployment Strategies
Recreate: Stop Everything, Then Start the New Thing
Rolling: Two Versions, One Database
Blue/Green: Paying for the Fastest Rollback There Is
Canary: One Percent, Then Five, Then Watch
Feature Flags: Deploy Is Not Release
Progressive Delivery: Exposure as a Dial
Version Coexistence: N and N+1, in Both Directions
Canary Analysis: Compared Against What?
A Successful Deploy Is Not Evidence of a Healthy System
Rollback: Only Useful If It Is Actually Safe
Roll Forward: When Going Back Is the Harder Option
Level 4

Infrastructure you can reproduce

The environment the artifact runs in becomes reviewable, reproducible and diffable rather than a console someone clicked. You can stand a whole environment up from a repository, see what a change will destroy before it destroys it, and detect when reality has stopped matching the code.

Infrastructure as Code
Declarative vs Imperative Infrastructure
The Plan: Desired vs Current
State
Drift
Destructive Changes: What a Rename Really Does
Immutable Infrastructure
Pets and Cattle, Read Carefully
Preview Environments
Ephemeral Environments
Manual Production Changes
Production Access
Least Privilege in Production
Level 5

Orchestration, discovery and traffic

You can run many replicas of many services, let them find each other, and put traffic in front of them without dropping requests. Kubernetes is taught here as one implementation of the scheduling and reconciliation problem — the level is passed by understanding the problem, not by adopting the tool.

Do You Need Kubernetes?
The Problems Kubernetes Answers
Cluster, Control Plane, Nodes, Pods
Pods: The Unit That Gets Scheduled
Deployments: Declaring What Should Be Running
ReplicaSets: The Layer You Should Not Manage
Services: A Stable Address Over Moving Pods
Getting Traffic Into the Cluster
Reconciliation: The Loop Under Everything
Apply Is Not Running
The Scheduler, and Why a Pod Is Pending
Requests and Limits
Probes: Readiness, Liveness and Startup
ConfigMaps and Secrets
Service Discovery in Operation
DNS in Production
Operating a Load Balancer
Draining: Stopping Without Dropping
Certificates as an Operational Object
Renewal: Automating the Thing That Expires
Level 6

Operating it, and responding when it breaks

You can tell whether production is healthy without asking anybody, get woken up only for things that need a human, and run an incident from detection to mitigation to a write-up that changes the system. This is the level that turns a deployer into an operator.

Using Observability, Not Building It
An Alert Should Demand Action
Alert on Symptoms, Not on Causes
Alert Fatigue
Dashboards an Operator Can Act On
Deploys on the Same Timeline as the Symptom
On-Call Is Production Ownership
Rotations People Can Sustain
What Happens Between the Page and the Postmortem
Severity: What It Should Reflect
Stop the Harm Before You Understand It
Roles During an Incident
Reconstructing What Actually Happened
Telling People What Is Happening
Postmortems
Root Cause vs Contributing Factors
Action Items That Change the System
Production Debugging
Deployment-Centric Debugging
Change Correlation
Level 7

Capacity, cost and the stateful parts

You can answer how much traffic the system can take, what happens when it takes more, what it costs per unit of work, and how to change a schema without an outage. This is where reliability, money and the database stop being separate conversations.

Capacity Management
Building a Capacity Model
Headroom
Load Shedding
Overprovisioning
Idle Capacity
Autoscaling
Choosing the Scaling Signal
How Autoscaling Fails
Horizontal Pod Autoscaling
Queue-Based Autoscaling
Scale to Zero
Cost Awareness
Cost Drivers
Cost Per Request
FinOps
Operating a Production Database
The Connection Budget
Why Migrations Are the Dangerous Change
Expand, Migrate, Contract
Zero-Downtime Migrations
A Migration and a Deploy Are One Event
Backfills
Destructive Migrations
Why Stateful Workloads Are Harder
Operating Queues and Scheduled Work
Operating a Cache
Level 8

Release engineering, supply chain and platform

You can answer "what is in production, where did it come from, and who approved it" with evidence rather than recollection — and you can make the safe path the easy path for other teams instead of reviewing their work by hand. This is the level where delivery becomes a product with users.

Release Engineering as a Discipline
The Release Manifest
Change Management
The Audit Trail
Promotion
Artifact Retention
The Delivery Chain as Attack Surface
Build Provenance
The Builder Is Inside the Trust Boundary
Signing and Verifying Artifacts
Software Bill of Materials
Scanning, and Why a Finding Is Not a Risk
Securing the Pipeline Itself
CI Security
Platform Engineering
The Internal Developer Platform
Golden Paths
Service Templates
Self-Service Infrastructure
Guardrails, Not Gates
Policy as Code
Developer Experience as an Operational Metric
How to Automate Something
The Automation Trap
Toil
Level 9

Multi-region, recovery and production engineering

You can lose a region, a database or a dependency and still have a procedure that returns the system to service within objectives you have actually tested. At this level you are no longer keeping a service up; you are engineering the properties that let an organisation keep many of them up.

Backup Operations
Restore Drills
Disaster Recovery as an Operation
RTO and RPO
Region Failover
Operating in More Than One Region
Capacity During Failover
Partial and Logical Data Recovery
Production Readiness Review
The Readiness Scorecard
The Ownership Record
Runbooks
Runbook Anti-Patterns
Break-Glass Access
Access Review
How Networks Fail in Production
Reducing Blast Radius
Production Anti-Patterns
CI/CD Anti-Patterns
Learning Across Incidents
Deploying an Agent
Prompts and Models Are Deployables
Canarying a Model or Prompt Change
The Agent Kill Switch
Agent Cost in Production