We Doubled the Cluster and It Got Slower
Read what each party saw, commit to a cause, and only then find out which of them was right. The root cause, why the obvious reading was wrong, and the fixes are all held back until you have answered.
The symptom
What was reported, before anyone knew what was happening.
What each party saw, and what each concluded
The evidence, in the form it actually arrived in — several parties, several partial views, several confident conclusions.
| Who | What they could actually see | What they concluded | Verdict |
|---|---|---|---|
| Shard 7 | 61% of all writes arriving at itself. | I am under-provisioned relative to my peers. | ✕ wrong |
| Capacity planning | Aggregate cluster CPU at 27%. Plenty of headroom. | The cluster is not the bottleneck; it must be a slow query. | ✕ wrong |
| The router | Hashing tenant_id and distributing uniformly across the ring, exactly as configured. | Distribution is correct and even. | ✓ right |
| On-call engineer | Adding 16 shards moved the hot key to shard 23 and left p99 higher. | The rebalance was botched. | ✕ wrong |
3 of 4 parties reasoned correctly from what they could see and still reached the wrong conclusion. Nobody in this table is careless. Each one acted on complete-looking local information, and the information was local. That gap — between what a node can observe and what is true — is the whole domain, and one of these readings will usually be yours.
Commit before you read on
The hash function is uniform and the router is behaving correctly. Commit before reading on: why does a uniform hash still produce a shard carrying 61% of writes, and why can no amount of extra shards fix it?
Write it down even if you are unsure. An unwritten guess quietly becomes “that is what I thought” the moment you read the answer.