What is On-Call Toil and How Do You Measure It?
On-call toil is the manual, repetitive, tactical, and automatable operational work an on-call engineer performs to keep production running, work that scales linearly with the size of the system and produces no lasting engineering value. It covers tasks like acknowledging alerts, hopping across dashboards to correlate signals, restarting services, and rebuilding incident context by hand. You measure it by tracking the share of on-call time spent on this work, most commonly targeting toil below 50% of an SRE team's time as popularized by Google SRE practice.
What counts as on-call toil?
On-call toil is work that is manual, repetitive, automatable, tactical, devoid of enduring value, and that grows as the system grows. If a task fits those markers, it is toil, whether or not it happens during an incident. Classic examples include acknowledging and triaging low-value alerts, restarting a stuck service, running the same diagnostic query for the tenth time, tool-hopping across metrics, logs, and traces to correlate a symptom to a cause, and re-investigating an incident the team already solved last month because nothing was captured.
On-call toil is any operational work that is manual, repetitive, automatable, and produces no lasting engineering value, and it scales linearly with the size of the system rather than with the value it delivers.
Toil is not the same as overhead like team meetings or performance reviews, and it is not the same as genuinely novel engineering work that requires human judgment. The distinguishing test is whether a machine could do the task if the right automation existed.
How do you measure on-call toil?
You measure on-call toil by quantifying the proportion of engineering time spent on manual operational work, then tracking that proportion over time. The most cited benchmark comes from Google SRE practice, which recommends keeping toil below 50% of an SRE's time so the remainder can go to engineering that reduces future toil. Beyond that headline ratio, teams instrument a set of concrete signals.
Toil that is not measured is toil that compounds silently, because work no one counts is work no one is funded to eliminate.
The practical method is a recurring survey plus telemetry: have on-call engineers log where their shift time went, correlate that with alert and incident tooling data, and review the trend every sprint or quarter. The table below compares the common metrics teams use to quantify on-call toil and what each one reveals.
| Metric | What it measures | Why it matters |
|---|---|---|
| Toil percentage | Share of on-call time spent on manual, repetitive work | The headline SRE metric; target is under 50% |
| Pages per shift | Alert volume reaching a human per on-call rotation | High counts signal alert fatigue and interrupt-driven work |
| Actionable alert rate | Percentage of alerts that require real human action | Low rates expose noise the team is manually filtering |
| Tools touched per incident | Distinct systems opened to investigate one incident | High counts reveal manual correlation toil |
| Time to root cause | Minutes spent identifying cause before any fix | Long times mean investigation, not repair, dominates |
| Repeat-incident rate | Share of incidents already solved before | High rates mean fixes and context are not being captured |
Why on-call toil matters to the business
Unmanaged on-call toil is a reliability risk and a roadmap tax, not just an engineer-wellbeing concern. When toil consumes an on-call rotation, alert fatigue sets in and the failures that actually matter start slipping through. Industry survey data underlines the scale of the drain: NeuBird AI's 2026 State of Production Reliability and AI Adoption Report found that 83% of teams navigate four or more tools during a live incident, and that 40% of engineering time goes to incident management rather than building product.
When toil rises, alert fatigue rises with it, and the incidents that matter most are the ones most likely to be tuned out.
Toil also concentrates in the few senior engineers who understand critical services, which means burnout lands precisely on the people the business can least afford to lose. Measuring toil turns an invisible drain into a number leaders can act on: it converts "our on-call is rough" into a tracked percentage that justifies investment in automation and prevention.
How do you reduce on-call toil?
You reduce on-call toil by attacking its sources: cut the noise that generates low-value pages, automate the repetitive investigation and remediation steps, and capture every fix so the same problem is never solved twice. The highest-leverage move is upstream: raising alert signal quality so fewer pages fire at all, rather than only responding to the existing queue faster.
NeuBird AI is a Production Ops Agent platform, a platform of specialized agents orchestrated as one, that addresses on-call toil across three pillars: Prevent (catch degradation before the page), Resolve (investigate and resolve incidents autonomously), and Operate (capture every fix and automate routine work between incidents). Rather than adding another dashboard for an engineer to read, it acts on the operational work directly, inside the customer's own environment with human-in-the-loop guardrails and a full audit trail.
The durable way to cut on-call toil is to change which pages happen, not just to answer the same pages faster.
See how NeuBird AI approaches on-call toil reduction and supports SRE teams reducing on-call load for a deeper treatment of each pillar, or read the evaluation guide for alert fatigue and noise reduction.
What to remember
- 1On-call toil is manual, repetitive, automatable operational work that scales with system size and produces no lasting engineering value.
- 2The headline benchmark from Google SRE practice is keeping toil below 50% of an SRE team's time.
- 3Measure toil with concrete signals: toil percentage, pages per shift, actionable alert rate, tools touched per incident, time to root cause, and repeat-incident rate.
- 4Unmanaged toil drives alert fatigue, burnout among senior engineers, and reliability risk, not just morale problems.
- 5The durable way to cut toil is upstream: raise alert signal quality so fewer pages fire, then automate investigation and capture every fix.
Frequently asked questions
What is the difference between on-call toil and operational overhead?
Toil is manual, repetitive, automatable work directly tied to running production that scales with system size, like restarting services or triaging alerts. Overhead is administrative work like meetings, reviews, and planning. The key test for toil is whether a machine could do the task if the right automation existed; overhead generally cannot be automated away.
What percentage of time should be spent on toil?
Google SRE practice recommends keeping toil below 50% of an SRE team's time so the remaining capacity can go to engineering work that reduces future toil. Above that threshold, teams tend to fall into a reactive cycle where operational load consumes the very hours needed to automate it away, and toil compounds.
How do you measure on-call toil in practice?
Combine a recurring survey with telemetry. Have on-call engineers log where their shift time went, then correlate that with data from alerting and incident tooling: pages per shift, actionable alert rate, tools touched per incident, time to root cause, and repeat-incident rate. Review the toil percentage trend every sprint or quarter to see whether it is rising or falling.
Does reducing alerts reduce on-call toil?
Reducing low-signal alerts is one of the highest-leverage ways to cut toil, because a large share of on-call work is triaging noise that never needed human action. Raising alert signal quality upstream so fewer pages fire attacks toil at its source, rather than only helping engineers respond to the existing alert queue faster.
See it in action. No slides.
NeuBird AI compresses incident investigation from hours to minutes: autonomous root cause analysis, with zero manual triage.