Refuse to start the incident-response conversation from the premise that an on-call system is needed - require an SLA first, and put accepting overnight downtime on the table as a real option
Situation
Ryan brought Peter a plan: after the Portal outage he had asked Justin and Nathan whether engineering had an on-call rotation, got a no, and wanted to build one at the Tuesday Aug 18 meeting. Peter agreed to the meeting, added Chris Wolford to the invite, and then inverted the agenda. His first question would be do we need this - not dismissively, but because he wanted both branches evaluated: accept that some systems are down overnight and publish a status page the way Apple does, or commit to short SLAs and staff an on-call rotation to meet them. He named the decision rule: define the SLA for each class of thing first, and if any SLA is under 12 hours, that implies on-call; if none is, it does not.
Reasoning
Peter separated the felt problem (Portal was down and nobody could fix it) from the actual question (what response time has CIQ committed to). On-call is a cost - it is a permanent tax on a team he had just told Bjorn is already overloaded - and he will not pay it to satisfy an intuition. His stated fear was the opposite failure too: I do not want things to stay down overnight because of inertia and because we did not have a conversation. So the SLA is not a delaying tactic, it is the artifact that makes either answer a decision rather than a default. He drew on running orgs at Apple where most teams had no on-call system despite customer-facing surfaces.
Additional Context
Follows the Portal outage and a separate four-week failure of the Zendesk-to-Slack paging integration that Steve surfaced the same day. Both incidents make the on-call answer feel obvious, which is precisely why Peter forced the criterion ahead of the conclusion.
Observed Evidence
Peter: one of the first questions I am going to ask, though, and I do not mean it dismissively, is do we need this. I have run piles of orgs at Apple. Most of them did not have an on-call system. And: that could be something we choose to accept as a company. ... We could do the same thing, or we could have teams committing to bringing it back up as quickly as possible. They both have pluses and minuses. And I want to go into the discussion evaluating both sides of that.
Matching Patterns
Confidence Breakdown
Reasoning Depth Analysis
People Involved
Source
reflection
AI Confidence
87%
Related Context
fathom
I want to put on the plate what I want to see out of it is an SLA for certain types of things. And then if we have got an SLA that is shorter than 12 hours, that means we need an on-call system. Or we do not have an SLA that is shorter than 12 hours.
Outcome
No outcome recorded yet.
Decision ID: 9ede202d-9eb5-4ecd-b1c0-48dd1b10fc41