<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>On-Call on David R. Longnecker - Converting Coffee to Code</title><link>https://drlongnecker.com/tags/on-call/</link><description>Recent content in On-Call on David R. Longnecker - Converting Coffee to Code</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sun, 06 Sep 2026 09:00:00 -0600</lastBuildDate><atom:link href="https://drlongnecker.com/tags/on-call/index.xml" rel="self" type="application/rss+xml"/><item><title>The Alert Budget</title><link>https://drlongnecker.com/blog/2026/09/the-alert-budget/</link><pubDate>Sun, 06 Sep 2026 09:00:00 -0600</pubDate><guid>https://drlongnecker.com/blog/2026/09/the-alert-budget/</guid><description>&lt;p&gt;Running technology teams across several industries taught me how much of &amp;ldquo;building products&amp;rdquo; is supporting them.&lt;/p&gt;
&lt;p&gt;Every on-call rotation has a ceiling. Past a certain number of incidents per shift, the team stops being able to recover between them. Everyone counts how many calls they take and how many issues they solve, but few track the recovery time between.&lt;/p&gt;
&lt;h2 id="google-ran-the-numbers"&gt;Google Ran the Numbers&lt;/h2&gt;
&lt;div class="stat-callout stat-callout-coral" role="figure" aria-label="Stat: 2. incidents per 12-hour shift before quality degrades 6 hours average to close one incident, root cause through postmortem. Source: Google SRE Book, Chapter 11, Being On-Call"&gt;
 &lt;div class="stat-callout-bar"&gt;&lt;/div&gt;
 &lt;div class="stat-callout-content"&gt;
 &lt;div class="stat-callout-number"&gt;2&lt;/div&gt;
 &lt;div class="stat-callout-label"&gt;incidents per 12-hour shift before quality degrades&lt;/div&gt;
 &lt;div class="stat-callout-label"&gt;6 hours average to close one incident, root cause through postmortem&lt;/div&gt;
 &lt;div class="stat-callout-source"&gt;Google SRE Book, Chapter 11, Being On-Call&lt;/div&gt;
 &lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;The Google SRE team published its version of this years ago and, in my experience, it still holds up. Their &lt;a href="https://sre.google/sre-book/being-on-call/"&gt;SRE Book chapter on being on-call&lt;/a&gt; (Chapter 11) found that fully closing out a single incident, from root-cause work through the postmortem, takes about six hours on average. Do that math against a 12-hour shift and the ceiling falls out. Two incidents is the most a shift can absorb before something gets rushed. Google states it as a target of fewer than two paging events per shift.&lt;/p&gt;</description></item></channel></rss>