Refuse to start the incident-response conversation from the premise that an on-call system is needed - require an SLA first, and put accepting overnight downtime on the table as a real option

August 13, 2026 at 10:42 PMoperationalmedium

Situation

Ryan brought Peter a plan: after the Portal outage he had asked Justin and Nathan whether engineering had an on-call rotation, got a no, and wanted to build one at the Tuesday Aug 18 meeting. Peter agreed to the meeting, added Chris Wolford to the invite, and then inverted the agenda. His first question would be do we need this - not dismissively, but because he wanted both branches evaluated: accept that some systems are down overnight and publish a status page the way Apple does, or commit to short SLAs and staff an on-call rotation to meet them. He named the decision rule: define the SLA for each class of thing first, and if any SLA is under 12 hours, that implies on-call; if none is, it does not.

Reasoning

Peter separated the felt problem (Portal was down and nobody could fix it) from the actual question (what response time has CIQ committed to). On-call is a cost - it is a permanent tax on a team he had just told Bjorn is already overloaded - and he will not pay it to satisfy an intuition. His stated fear was the opposite failure too: I do not want things to stay down overnight because of inertia and because we did not have a conversation. So the SLA is not a delaying tactic, it is the artifact that makes either answer a decision rather than a default. He drew on running orgs at Apple where most teams had no on-call system despite customer-facing surfaces.

Additional Context

Follows the Portal outage and a separate four-week failure of the Zendesk-to-Slack paging integration that Steve surfaced the same day. Both incidents make the on-call answer feel obvious, which is precisely why Peter forced the criterion ahead of the conclusion.

Observed Evidence

Peter: one of the first questions I am going to ask, though, and I do not mean it dismissively, is do we need this. I have run piles of orgs at Apple. Most of them did not have an on-call system. And: that could be something we choose to accept as a company. ... We could do the same thing, or we could have teams committing to bringing it back up as quickly as possible. They both have pluses and minuses. And I want to go into the discussion evaluating both sides of that.

Matching Patterns

74%
Purpose Is the Decision Procedure(refuses to engage the proposal on its merits, restates what the meeting must produce - an SLA - and derives the answer from it)
66%
Reclassify to Route - Decide Who Owns It, Then Give a Criterion Not a Verdict(supplies a criterion (12-hour SLA threshold) rather than a verdict on on-call)
52%
Protect Engineering Capacity(declines to add a standing load to an already-overloaded org without justification)

Confidence Breakdown

34/35
Evidence
27/30
Pattern
19/20
Source
7/15
Corroboration

Reasoning Depth Analysis

Org Signal:Signals that a visible incident does not automatically buy a permanent process. The org is expected to price its commitments before making them.
Who Affected:Every engineer who would carry a pager, plus Support and customers who would be governed by whatever SLA is written.
Precedent:Strong one: post-incident proposals arrive at Peter needing a stated commitment level, not a stated remedy. Reusable for any future outage-driven process ask.
Consequences:Real. If the SLAs come back long, there will be no on-call rotation, and Peter has pre-committed to accepting overnight downtime publicly rather than quietly.
Timing:Now rather than later because Ryan had already convened the meeting; framing it before Tuesday is the only way to stop the meeting from ratifying a foregone conclusion.

Related Context

🎥
Ryan <> Peter Weekly 1:1

fathom

I want to put on the plate what I want to see out of it is an SLA for certain types of things. And then if we have got an SLA that is shorter than 12 hours, that means we need an on-call system. Or we do not have an SLA that is shorter than 12 hours.

Outcome

No outcome recorded yet.

Decision ID: 9ede202d-9eb5-4ecd-b1c0-48dd1b10fc41