Rotations People Can Sustain
Load, frequency and recovery are properties of the system, and a rotation that cannot be sustained is a defect in the system rather than a shortcoming of a person.
The question, the obvious approach, and why it breaks
Every lesson starts where the work starts: an operational problem, a first attempt that is entirely reasonable, and the way production disagrees with it.
What makes an on-call rotation something a team can carry for years rather than months?
Being reachable outside working hours has a real and cumulative cost, and a rotation that exceeds what people can absorb degrades response quality first and retains people second.
On-call is part of the job. People manage. If someone is struggling, they can raise it.
"People manage" until they do not, and the failure is discretely visible only at resignation — long after response quality started degrading.
- "People manage" until they do not, and the failure is discretely visible only at resignation — long after response quality started degrading.
- Raising it individually asks a person to describe a shared, structural problem as a personal difficulty, which is a bad position to be in and reliably delays the conversation.
- Sleep loss degrades exactly the judgement an incident needs, so an overloaded rotation makes its own incidents longer, which makes the load worse.
- The people who leave first are often those with the least slack outside work — caring responsibilities, health conditions — which narrows the team in ways nobody chose.
- The cost is invisible in every dashboard the organisation looks at. Nothing reports "the rotation is not sustainable" until it stops working.
What is actually happening
Underneath the tooling, which is the part that survives a change of tool.
- Three things determine sustainability, and they are not the same. Frequency: how often a person is on call. Load: how many interruptions per shift, and how many outside sleeping hours. Recovery: whether there is real time to recover afterwards.
- They interact. A frequent rotation with a quiet pager is very different from an infrequent one where every shift is broken sleep, and averaging pages per week hides both.
- Frequency is a staffing arithmetic. If a rotation is 24/7 with weekly shifts, a team of four means one week in four; a team of eight means one in eight. Below a certain number of people, no schedule makes it sustainable — the honest answer is more people, a smaller coverage window, or fewer services, not a cleverer rota.
- Load is a property of the alert set and of the system's reliability, both of which are changeable. Most unsustainable rotations are unsustainable because of pages that should not exist (Alert Fatigue).
- The most important framing: an unsustainable rotation is a system problem. It reflects the number of people, the reliability of the services, the quality of the alerting and the absence of time to fix things. It is not evidence about the resilience of any individual carrying it, and treating it as such prevents the only fixes that work.
- Response quality degrades before people leave. Slower acknowledgement, more escalations, and more mitigation-by-restart are the leading indicators, and they show up in the incident record before anyone resigns.
Three variables, and what actually moves each
Teams tend to discuss the rota when the problem is the alert set, or discuss the alerts when the problem is that there are four people covering a service that needs eight.
Separating the three makes the conversation tractable, and makes it about the system rather than about who is coping.
| Variable | What it is | What moves it | Signal it is the binding constraint |
|---|---|---|---|
| Frequency | How often a person is on call | Number of people; width of the coverage window | Shifts feel relentless even when they are quiet |
| Load | Interruptions per shift, especially at night | Alert quality; reliability of the services covered | Shifts are rare and every one costs sleep (Alert Fatigue) |
| Recovery | Real time to recover after a disturbed shift | Team norms and whether the plan has slack | People return to full delivery load the next morning |
| Scope | How many services one rotation covers | Ownership boundaries; onboarding of new services | Pages arrive for systems the responder cannot act on |
| Depth | How many people can genuinely handle a page | Runbooks, shadowing, rotation of the coordinator role | One person is escalated to every time (Runbooks) |
What a rotation review asks
Held on a schedule, with the team, about the rotation rather than about anyone in it. The framing matters: these are questions about a system that the team can change, not questions about how well people are holding up.
The last step is the one that gets dropped, and dropping it is what turns a review into a discussion.
- 1Read the numbers
Pages per shift, out-of-hours pages, duration, distribution across people.
fails by Using a quarterly average, which hides the bad shifts entirely.
evidence Per-shift figures, not aggregates.
- 2Find the concentration
Identifies which alert or component produced most of the load.
fails by Treating load as diffuse when it is nearly always concentrated.
evidence A ranked list; usually the top one or two dominate.
- 3Check depth
Asks how many people could genuinely have handled each page.
fails by Counting the rota rather than the capability.
evidence Escalation pattern — repeated escalation to one person is the finding.
- 4Check recovery
Asks whether people took real time after disturbed nights.
fails by Asking whether it was available rather than whether it was taken.
evidence It actually happened, and nobody had to request it as a favour.
- 5Decide one structural change
Picks the single highest-load source and commits capacity to removing it.
fails by Producing a list of intentions with no owner.
evidence Named owner, allocated capacity, in the current plan (Action Items That Change the System).
- 6Re-check next cycle
Compares the numbers against last time.
fails by Never happening, which is how load ratchets upward unnoticed.
evidence A trend, not a snapshot.
If the review consistently concludes that the numbers are fine and people say otherwise, believe the people and look for what the numbers are not capturing — a single recurring 4am page is worse than its count suggests.
How a rotation degrades
The path below is gradual and each step is individually tolerable, which is exactly why it runs to completion. Nothing in it is about anyone's capability.
It is also reversible at every point, and cheaper to reverse the earlier you notice. The leading indicator is response quality, not how anyone says they are feeling.
How to do it properly
Most important first.
- Measure the load honestly: pages per shift, how many outside working hours, how many broke sleep, and how long each took. Averages across a quarter hide the shifts that did the damage.
- Fund the fix. Reserve explicit capacity each cycle for reducing the causes of pages, or the rotation will consume the time that would have improved it (Toil).
- Give recovery time after a disturbed night as a stated norm rather than a favour someone has to ask for. Asking is the part that fails.
- Set a load threshold that triggers action rather than discussion — a shift above it means the top page source gets fixed before feature work, by default.
- Have a named secondary so nobody is the only person able to respond, and so being unavailable for an hour is possible.
- Fix the biggest page source every cycle. Load is nearly always concentrated in a few alerts or one unreliable component.
- Review the rotation as a system with the team, on a schedule, separately from any individual's experience of it — the point is to make it discussable without anyone having to volunteer that they are struggling.
- Compensate it, in time or money, according to what your organisation and jurisdiction allow. Unpaid, unacknowledged availability is where resentment starts.
How much can this affect
Every production change has a blast radius. Stated as a scale so it is comparable between changes rather than adjectival — and paired with what actually contains it, because a wide scope with a real containment mechanism is a different situation from a wide scope with none.
Nothing contains it — a degraded rotation lengthens every incident across every service it covers, and the degradation is invisible until it is attrition.
What can go wrong
- Load measured as an average, which hides that one person had four bad nights and everyone else had none.
- Improvement work planned and continuously displaced by delivery pressure, so the rotation never improves.
- A rotation carried in practice by one person who is always available, whose absence would be an operational crisis.
- Recovery time nominally available and culturally impossible to take.
- Adding people to a rotation covering services they cannot act on, which reduces frequency and increases escalations.
- Treating fatigue as an individual matter — coaching the person rather than changing the load (Root Cause vs Contributing Factors).
- Growth in services outpacing growth in the rotation, so load rises without any decision being made.
- "Some people just handle on-call better." Some people currently have more slack in their lives. Designing a rotation around that selects your team for personal circumstances rather than for skill.
- "Burnout is a personal resilience issue." It is a load issue. The variables that determine it — number of people, number of pages, time to fix things — are all organisational.
- "We only had three pages last month, so the rotation is fine." Three pages at 4am on the same person's shift is a different fact from three spread across a quarter.
- "Nobody has complained." The people most affected are usually the least able to complain, and the complaint arrives as a resignation.
- "Compensation makes it sustainable." Compensation makes it fair. Load is what makes it sustainable, and paying more for an unsustainable rotation buys a little time and no improvement.
Operating it
Evidence is the signal, not the intention. Rollback is sometimes 'you cannot, and that is the point'.
- Pages per shift, and out-of-hours pages per shift, are tracked and trending down.
- Most shifts pass with no out-of-hours interruption.
- Time spent reducing page sources appears in the plan and actually happens.
- People take recovery time after disturbed nights without asking permission.
- Nobody on the team is structurally irreplaceable on the rotation.
- If a rotation is not sustainable, the available moves are all structural: reduce the coverage window, reduce the services covered, add people, or fix the top page sources. Asking people to endure it is not one of them, and choosing it is choosing the attrition.
- Removing a service from a rotation temporarily — accepting slower response for it while its reliability is fixed — is a legitimate and often correct decision, and it should be made explicitly rather than by exhaustion.
- Automate the measurement: pages per shift, time of day, duration, and which alert generated them. Nobody assembles this manually, and without it the conversation is anecdotal.
- Automate schedule fairness — even distribution of nights, weekends and holidays — and make overrides easy so being unavailable is a normal thing rather than a negotiation.
- Automate the removal of pages: every automated safe remediation is a night someone sleeps (How to Automate Something).
- Do not automate the assessment of whether the rotation is sustainable. Numbers inform it; the people carrying it are the ones who know, and the mechanism to ask them has to be a routine, not an escalation.
- Capacity spent reducing pages is capacity not spent on features, and the return arrives as absence — nights that were quiet, incidents that did not happen.
- More people on a rotation lowers frequency and raises the number who must stay familiar with the systems; familiarity has its own maintenance cost.
- Narrower coverage windows are kinder and mean some failures wait until morning. That is a legitimate trade when the service permits it, and it should be an explicit decision.
Where this applies
This domain is unusually tool- and organisation-dependent. These labels say what each claim is specific to, and what a different platform, provider or organisation does instead.
- ORG-SPECIFICCompensation, working-time limits, whether out-of-hours availability is contractual, and what recovery time is permitted differ by employer and by jurisdiction — in some, rest periods after night work are statutory rather than discretionary. What generalises is the relationship between frequency, load and recovery; the entitlements do not, and they must be checked locally.
- SCALE-SPECIFICSustainability is arithmetic before it is culture. A 24/7 rotation across a team of three cannot be made sustainable by scheduling, however it is arranged; the fixes are more people, a narrower window or fewer services. Above roughly six or eight people, load rather than frequency becomes the binding constraint.
Where the depth lives
This domain teaches delivery and operation, and hands the mechanism off to the domain that owns it.
- — Testing & Reliability Engineering — treating operational load as a measured property of the system, in the same way reliability is.