Distributed Systems Roadmap
From what actually changes when a call crosses a machine boundary, through replication, consistency, time, consensus and delivery, to running a distributed system in production and knowing when not to build one at all. Ten levels, every lesson in the domain placed exactly once.
- 1
Remote Calls and Partial Failure
0 / 18 masteredStop treating a network call as a function call. Everything else in the domain depends on this shift.
What Actually Makes a System DistributedFundamentalsA Remote Call Is Not a Function CallFundamentalsThe Network Changes EverythingFundamentalsNo Shared Memory: Every Node Sees a CopyFundamentalsThere Is No Global ClockFundamentalsWhy Distribute At AllFundamentalsWhen Not to DistributeFundamentalsWhere the Boundary GoesFundamentalsName the Invariant Before You Choose the ProtocolFundamentalsPartial Failure: The Founding ConditionFailure ModelsA Timeout Tells You Nothing About Whether It HappenedFailure ModelsFailure Models: What You Are Allowed to AssumeFailure ModelsNo Heartbeat Does Not Mean DeadFailure ModelsCrashed or Just Slow: The Distinction You Cannot MakeFailure ModelsByzantine Failures, and Why You Probably Do Not Assume ThemFailure ModelsFault Domains: What Fails TogetherFailure ModelsCorrelated Failure: The Independence Assumption Is Usually FalseFailure ModelsRedundancy Is Not ResilienceFailure Models - 2
Replication and Consistency
0 / 18 masteredCopies buy availability and cost you agreement. Learn to name the guarantee precisely.
Why Replicate: What a Second Copy Buys YouReplicationLeader-Based Replication: Buying Order With a Single WriterReplicationSynchronous Replication: Paying Latency for a Durability GuaranteeReplicationAsynchronous Replication: The Loss Window You ChoseReplicationMulti-Leader Replication: Accepting Writes in More Than One PlaceReplicationRead-After-Write: Letting a User See Their Own ChangeReplicationMonotonic Reads: Never Let Time Run BackwardsReplicationQuorums: What R + W > N Does and Does Not BuyReplicationLeaderless Replication: Every Replica Accepts WritesReplicationConsistency Models: What "Consistent" Has To MeanConsistency ModelsLinearizability: An Operation Is an Interval, Not a PointConsistency ModelsSerializability vs Linearizability: Two Different PropertiesConsistency ModelsEventual Consistency: If Updates Stop, Replicas ConvergeConsistency ModelsCausal Consistency: Never Show an Effect Before Its CauseConsistency ModelsSession Guarantees: The Underrated Middle GroundConsistency ModelsCAP: What the Theorem Actually SaysConsistency ModelsPACELC: The Trade-off That Exists When Nothing Is BrokenConsistency ModelsChoosing a Consistency Model: Start From the InvariantConsistency Models - 3
Time, Ordering and Conflict
0 / 14 masteredClocks disagree. Causality is what you can actually rely on — and when two writes are concurrent, someone must decide.
Two Timestamps Are Not an OrderingTime & OrderingClock Skew: The Gap You Cannot Measure From InsideTime & OrderingNever Measure a Duration With the Wall ClockTime & OrderingHappens-Before: The Only Ordering You Actually HaveTime & OrderingLamport Clocks: Consistent With Causality, Blind to ConcurrencyTime & OrderingVector Clocks: Buying Concurrency Detection at O(N)Time & OrderingFour Orderings, Four PricesTime & OrderingTotal Order Broadcast Is Consensus Wearing a Different HatTime & OrderingTwo Writes, No Order, One Answer RequiredConflict ResolutionLast Write Wins Is Data Loss You Chose by DefaultConflict ResolutionVersion Vectors: Making the Conflict VisibleConflict ResolutionOnly the Application Knows What the Merge MeansConflict ResolutionCRDTs: Deterministic Merge, Not Correct MergeConflict ResolutionWhat "Eventually Converges" Actually RequiresConflict Resolution - 4
Consensus and Coordination
0 / 18 masteredAgreement under failure, what it assumes, what it costs, and how often you can avoid needing it.
What Consensus Actually SolvesConsensusConsensus Is Not Magic: The Assumptions It Runs OnConsensusLeader Election: Choosing One, and Knowing You Were ChosenConsensusTerms and Epochs: Making Stale Leaders HarmlessConsensusSplit-Brain: Two Nodes, Both Certain They Are In ChargeConsensusFencing Tokens: Making the Stale Actor Safe, Not Just UnlikelyConsensusRaft: Elections, Terms and Three StatesConsensusThe Raft Log: Commit Index, Divergence and ReconciliationConsensusPaxos and the Other Protocols: What They Share and Where They DifferConsensusDo You Actually Need Consensus?ConsensusCoordination Couples AvailabilityCoordinationCoordination Avoidance: Restructuring the Problem Instead of Paying for ItCoordinationStart From the Invariant, Not From the ArchitectureCoordinationDistributed Locks: What They Are Actually ForCoordinationLeases: Authority With an Expiry DateCoordinationThe Stale Lock Holder: A Paused Process Does Not Know It Was PausedCoordinationCoordination Services: The Primitives, Not the ProductCoordinationDistributed Uniqueness: One Name, Many ShardsCoordination - 5
Messaging and Streams
0 / 16 masteredBrokers, queues and logs — and the delivery guarantee each one actually provides.
What a Broker Actually Buys YouMessagingWork Queues: One Task, One Worker, Competing ConsumersMessagingPublish/Subscribe: One Event, Many Independent ReadersMessagingQueue or Pub/Sub: Answer the Question in One SentenceMessagingAcknowledgement: The Two-Line Protocol That Decides Your Delivery SemanticsMessagingVisibility Timeout: The Message Is Hidden, Not YoursMessagingPoison Messages: The One That Fails Every Time, ForeverMessagingA Dead-Letter Queue Is a Workflow, Not a BinMessagingThe Log Is Not a QueueMessagingA Topic Is Not One Log: Ordering Lives Inside a PartitionStream ProcessingConsumer Groups: Queue Semantics Inside, Pub/Sub Semantics AcrossStream ProcessingRebalancing: Everyone Stops So the Partitions Can MoveStream ProcessingCommit Before or After: There Is No Third OptionStream ProcessingTwo Clocks: When It Happened and When You Saw ItStream ProcessingLate Events: The Window Already FiredStream ProcessingWatermarks: A Guess About Time, Made Precise Enough to Act OnStream Processing - 6
Partitioning and Membership
0 / 14 masteredSplitting data across nodes, and knowing which nodes are still there.
Why Partition: Four Ceilings, Four Different AnswersPartitioning & ShardingHash Partitioning and the Modulo TrapPartitioning & ShardingRange Partitioning: Scans You Keep, Hotspots You InheritPartitioning & ShardingThe Ring: Keeping the Mapping Stable When Membership ChangesPartitioning & ShardingVirtual Nodes: Many Positions per Machine, and Why It Is Not OptionalPartitioning & ShardingHot Partitions: The Skew Hashing Cannot FixPartitioning & ShardingRebalancing: A Load Spike You Schedule for YourselfPartitioning & ShardingCross-Partition Operations: Paying for What the Split Took AwayPartitioning & ShardingDiscovering Services: The Registry Is a Distributed System TooMembership & DiscoveryCluster Membership: A Belief, Not a FactMembership & DiscoveryGossip: Epidemic Spread Instead of Everyone Telling EveryoneMembership & DiscoveryAnti-Entropy: Repairing Divergence Nobody ReportedMembership & DiscoveryMerkle Trees: Finding the Difference Without Reading the DataMembership & DiscoveryFrom Alive-or-Dead to a Suspicion LevelMembership & Discovery - 7
Workflows and Idempotency
0 / 13 masteredAtomicity across services you do not control, and making retries safe.
Atomicity Stops at the Process BoundaryDistributed Transactions & SagasTwo-Phase Commit: Buying Atomicity With a PromiseDistributed Transactions & SagasThe Blocking Window: When 2PC Stops and WaitsDistributed Transactions & SagasSagas: Trading Isolation for AvailabilityDistributed Transactions & SagasA Refund Is Not a RollbackDistributed Transactions & SagasOrchestration: One Component Owns the WorkflowDistributed Transactions & SagasChoreography: The Workflow Nobody Wrote DownDistributed Transactions & SagasThe Retry Is a Decision, Not a ReflexIdempotency & DeliveryIdempotent Is a Property of the Whole Effect, Not the WriteIdempotency & DeliveryWhat Counts as the Same Operation?Idempotency & DeliveryWhere You Put the Acknowledgement Decides EverythingIdempotency & DeliveryExactly-Once Is a Scope, Not a GuaranteeIdempotency & DeliveryDeduplication: Bounded Memory Against an Unbounded StreamIdempotency & Delivery - 8
Overload, Deadlines and Caching
0 / 18 masteredBounded behaviour when demand exceeds capacity, and time budgets across a call graph.
Backpressure Is a Signal That Has to Travel — and Reach Someone Who Can Slow DownOverload & BackpressureRejecting Work on Purpose — and Rejecting It Cheaply Enough to HelpOverload & BackpressureDecide at the Door Whether the Capacity ExistsOverload & BackpressureOne Retry per Tier Is Not One Retry — It MultipliesOverload & BackpressureCap Retries as a Fraction of Traffic, Not as a Count per RequestOverload & BackpressureWithout Jitter, Every Client That Failed Together Retries TogetherOverload & BackpressureContainment Is Decided by What Is Shared, Not by Where the Service Boundaries AreOverload & BackpressureBulkheads: Buying Independence by Giving Up UtilisationOverload & BackpressureA Deadline Is Divided Across the Call Chain, Not Repeated at Every HopDeadlines & Tail LatencyPass the Remaining Budget Down, Not a Fresh OneDeadlines & Tail LatencyThe Caller Is Gone — Stopping Is Usually Right and Sometimes UnsafeDeadlines & Tail LatencySend a Second Request After p95 and Take Whichever Answers FirstDeadlines & Tail LatencyFan Out to 100 and the Component’s Tail Becomes the System’s MedianDeadlines & Tail LatencyA Cache Across Machines Is a Replica With No Replication ProtocolDistributed CachingInvalidation Is a Messaging Problem, Which Is Why Cache Bugs Are HardDistributed CachingOne Key Expires and Five Hundred Instances Miss at the Same MillisecondDistributed CachingYou Cannot Enumerate the Caches, So TTL Is the Bound and Invalidation Is the OptimisationDistributed CachingSharding Does Not Help a Single KeyDistributed Caching - 9
Storage, Compute and Geography
0 / 19 masteredThe layers beneath a distributed database, and what physics does to a multi-region design.
A Distributed Database Is a Stack, Not a BoxDistributed StorageDistributed File Systems: Chunks, a Metadata Service, and Where the Copies GoDistributed StorageObject Storage: A Flat Namespace With a Hash Behind ItDistributed StorageAcknowledged, Durable, Replicated: Three Different ThingsDistributed StorageRecovered State Is a Checkpoint Plus the Log After ItDistributed StorageA Consistent Cut, Without Stopping the WorldDistributed StorageSplitting a Computation Across MachinesDistributed ComputeMapReduce: The Model That Made the Trade-offs VisibleDistributed ComputeThe Shuffle Is the JobDistributed ComputeMove the Computation to the DataDistributed ComputeWho Runs What, and What Happens When a Worker Goes QuietDistributed ComputeOne Slow Task Sets the Pace for EverythingDistributed ComputeA Region Boundary Is a Consistency DecisionMulti-Region SystemsThe One Number You Cannot OptimiseMulti-Region SystemsThree Ways to Accept a Write in More Than One PlaceMulti-Region SystemsActive-Passive: Simple to Reason About, Rarely TestedMulti-Region SystemsActive-Active: Every Conflict Scenario Becomes RealMulti-Region SystemsEU and US Are Partitioned. Can Both Keep Accepting Writes?Multi-Region SystemsWhen the Data Is Not Allowed to LeaveMulti-Region Systems - 10
Operating Distributed Systems
0 / 22 masteredWhere to cut the system, how it fails in production, and how it recovers.
Detect, Contain, Recover, Reconcile, VerifyFailure & Recovery in ProductionGraceful Degradation: Which Dependency Is Actually CriticalFailure & Recovery in ProductionThe Steady-State Hypothesis and the Abort ConditionFailure & Recovery in ProductionCascading Failure: When the Response to Failure Causes More FailureFailure & Recovery in ProductionDependency Blast Radius: What Dies If This Node DiesFailure & Recovery in ProductionChaos Engineering Is Not Randomly Breaking ProductionFailure & Recovery in ProductionFault Injection: The Catalogue, and Which Faults Are HardFailure & Recovery in ProductionDistributed Debugging: The Question LadderFailure & Recovery in ProductionThree Nodes, Three Logs, and You Cannot Sort by TimestampFailure & Recovery in ProductionMicroservices Are a Distribution Decision, Not a Scaling TechniqueDistribution BoundariesThe Distributed Monolith: All of the Cost, None of the AutonomyDistribution BoundariesFour Questions That Test a Proposed BoundaryDistribution BoundariesThe Shared Database: An Honest Trade, Not a ProhibitionDistribution BoundariesExactly One Component Owns Each Piece of StateDistribution BoundariesSource of Truth: The Question Every Inconsistency Incident Is Really AskingDistribution BoundariesMaterialized Views: A Read Model That LagsDistribution BoundariesReconciliation Is a Component, Not a Cleanup ScriptDistribution BoundariesAn Agent System Is a Distributed SystemAgentic Distributed SystemsThe Model Retries Because It Cannot See the ResultAgentic Distributed SystemsAgents Do Not Negotiate. Processes Contend for State.Agentic Distributed SystemsResuming a Workflow That Died Halfway ThroughAgentic Distributed SystemsThe Failures That Produce No ErrorsAgentic Distributed Systems