From MTTR to Business Resilience
Beyond traditional metrics: how to measure and improve the true resilience of your engineering organization.
Beyond traditional metrics: how to measure and improve the true resilience of your engineering organization.
For decades, incident management has focused on a simple metric: Mean Time To Recovery (MTTR). It's straightforward to measure, easy to understand, and provides a clear target for improvement. Lower MTTR looks good on dashboards and in executive reports.
But MTTR tells only part of the story. An organization that resolves incidents in 30 minutes is not necessarily more resilient than one that takes 2 hours—especially if the first organization has 10 incidents per week while the second has one per month. Similarly, MTTR doesn't capture severity, doesn't account for data loss or customer impact, and ignores the systemic factors that lead to incidents.
True business resilience requires looking beyond individual incident metrics to understand how your organization responds to, learns from, and prevents disruptions.
Real resilience is multidimensional. It encompasses several interconnected capabilities:
How quickly does your organization notice when something goes wrong? Modern systems are complex, and problems can cascade silently. Some of the worst outages occur not because something broke, but because nobody knew it was broken.
Effective detection isn't just about monitoring tools. It requires:
Once detected, how quickly and effectively does your organization respond? MTTR is part of this, but not all of it. Response capability includes:
Can your team actually restore service? This requires more than knowing what to do. It requires:
What does your organization do after an incident? This is where most teams fall short. True resilience organizations:
If MTTR alone is insufficient, what metrics should replace or supplement it? Consider a balanced scorecard approach:
Change failure rate: What percentage of deployments cause incidents? Lower is better. Track deployment frequency alongside this to see if you're moving fast safely.
Mean time between failures (MTBF): How long does your system run without incident? This shifts the focus from fixing problems quickly to preventing them entirely.
Detection latency: How long between incident occurrence and detection? Measure this separately from resolution time. Fast detection allows faster response.
Detection accuracy: What percentage of alerts are valid incidents vs. false positives? High false-positive rates desensitize teams and reduce alertness.
Response time: How long between detection and first meaningful action? This measures organizational agility.
Escalation time: How long before the right people are engaged? Slow escalation extends outages unnecessarily.
Mean time to recovery (MTTR): Now in proper context—not as the primary metric, but as one component of a larger picture.
Recovery success rate: What percentage of recovery attempts succeed on the first try? Rollbacks that fail twice are particularly damaging.
Customer impact ratio: What percentage of users actually experience impact from a given incident? A backend service down for 30 minutes might affect 0% of users if it's not on the critical path.
Revenue impact: How much revenue was at risk during the outage? This bridges technical metrics to business outcomes.
Postmortem completion rate: What percentage of incidents have postmortems completed? Target 100% for any incident affecting customers.
Action item completion rate: What percentage of postmortem action items are completed? Track this monthly. Consistently low completion rates indicate either unrealistic action items or organizational commitment problems.
Recurrence rate: What percentage of incidents are repeats of previous incidents? High rates indicate learning isn't happening effectively.
Resilience isn't built through metrics alone. It requires systemic changes:
Invest in detection and observability: Spend on tools, training, and time to understand what's happening in your systems. Prevention is cheaper than recovery.
Build response muscle: Regular incident simulations and gamedays help teams execute under pressure. Like any skill, response capability improves with practice.
Invest in reliability engineering: Dedicate engineers to reducing system complexity, improving automation, and building resilience into architecture. This is not someone's part-time job.
Create a learning culture: Blameless postmortems only work if people believe they're actually blameless. Leaders must demonstrate this through action. Never punish learning. Instead, celebrate organizations that surface problems quickly.
Measure and monitor resilience: Use the metrics framework above. Review them monthly. Use them to guide investment and prioritization.
MTTR is a tactic; resilience is a strategy. While it's important to recover quickly from incidents, the real goal is to build organizations that prevent incidents, detect problems early, respond effectively, and learn continuously. When you focus on these broader dimensions, resilience becomes a competitive advantage that protects revenue, customer satisfaction, and team morale.
The most resilient organizations aren't those that fix things fastest—they're the ones that break things least often and learn most effectively.
The SRE Leadership Hub provides practical insights, frameworks, and advisory guidance for building reliable, scalable, and intelligent engineering organizations.
Explore how forward-thinking organizations are shifting reliability from an operational concern to a core business strategy.
Learn how to define meaningful SLOs and use error budgets as a powerful tool for balancing reliability and velocity.
Subscribe to our newsletter for engineering leadership perspectives and insights.