The Backup That Was on the Thing That Failed
Read what each party saw, commit to a cause, and only then find out which of them was right. The root cause, why the obvious reading was wrong, and the fixes are all held back until you have answered.
The symptom
What was reported, before anyone knew what was happening.
What each party saw, and what each concluded
The evidence, in the form it actually arrived in — several parties, several partial views, several confident conclusions.
| Who | What they could actually see | What they concluded | Verdict |
|---|---|---|---|
| Recovery process | The latest checkpoint on the replacement node dated 40 minutes before the failure, and a WAL segment directory that was empty. | Replay from the checkpoint; there is nothing else to apply. | ✓ right |
| Checkpoint job | A successful checkpoint every 40 minutes, written and fsynced, exit code 0. | A durable checkpoint exists. | ✕ wrong |
| Runbook test | A simulated node failure — a process kill — recovering in 4 minutes with no loss. | The recovery path works. | ✕ wrong |
| Storage engineer | The node’s replication factor of 3 and a healthy cluster. | Single-device loss is fully covered. | ✕ wrong |
3 of 4 parties reasoned correctly from what they could see and still reached the wrong conclusion. Nobody in this table is careless. Each one acted on complete-looking local information, and the information was local. That gap — between what a node can observe and what is true — is the whole domain, and one of these readings will usually be yours.
Commit before you read on
The checkpoints succeeded, the WAL was fsynced, and the cluster kept three replicas. Commit before reading on: name the fault domain assumption that every one of those three facts quietly depended on.
Write it down even if you are unsure. An unwritten guess quietly becomes “that is what I thought” the moment you read the answer.