Skip to main content

From MTTR to Business Resilience

Beyond traditional metrics: how to measure and improve the true resilience of your engineering organization.

Business Resilience Metrics

The Limitations of Traditional Metrics

For decades, incident management has focused on a simple metric: Mean Time To Recovery (MTTR). It's straightforward to measure, easy to understand, and provides a clear target for improvement. Lower MTTR looks good on dashboards and in executive reports.

But MTTR tells only part of the story. An organization that resolves incidents in 30 minutes is not necessarily more resilient than one that takes 2 hours—especially if the first organization has 10 incidents per week while the second has one per month. Similarly, MTTR doesn't capture severity, doesn't account for data loss or customer impact, and ignores the systemic factors that lead to incidents.

True business resilience requires looking beyond individual incident metrics to understand how your organization responds to, learns from, and prevents disruptions.

The Dimensions of Business Resilience

Real resilience is multidimensional. It encompasses several interconnected capabilities:

Detection Capability

How quickly does your organization notice when something goes wrong? Modern systems are complex, and problems can cascade silently. Some of the worst outages occur not because something broke, but because nobody knew it was broken.

Effective detection isn't just about monitoring tools. It requires:

  • Customer-focused monitoring that catches user-visible issues before internal alerts
  • Observability practices that surface anomalies and unexpected behavior
  • Clear escalation paths so detection leads to response
  • Regular testing of monitoring and alerting systems

Response Capability

Once detected, how quickly and effectively does your organization respond? MTTR is part of this, but not all of it. Response capability includes:

  • Clear incident command processes and role definitions
  • On-call rotations and escalation procedures
  • Rapid access to necessary tools, dashboards, and runbooks
  • Communication protocols with stakeholders
  • Authority to make emergency decisions without bureaucratic delays

Recovery Capability

Can your team actually restore service? This requires more than knowing what to do. It requires:

  • Redundancy and failover mechanisms built into architecture
  • Rollback capabilities for deployments
  • Tested disaster recovery procedures
  • Geographic distribution and multi-region capabilities where appropriate
  • Data backup and restore capabilities

Learning Capability

What does your organization do after an incident? This is where most teams fall short. True resilience organizations:

  • Conduct blameless postmortems within 24 hours
  • Identify root causes and systemic issues
  • Create action items and track them to completion
  • Share learnings across teams and organizations
  • Update runbooks, documentation, and monitoring based on learnings

Beyond MTTR: Measuring Organizational Resilience

If MTTR alone is insufficient, what metrics should replace or supplement it? Consider a balanced scorecard approach:

Prevention Metrics

Change failure rate: What percentage of deployments cause incidents? Lower is better. Track deployment frequency alongside this to see if you're moving fast safely.

Mean time between failures (MTBF): How long does your system run without incident? This shifts the focus from fixing problems quickly to preventing them entirely.

Detection Metrics

Detection latency: How long between incident occurrence and detection? Measure this separately from resolution time. Fast detection allows faster response.

Detection accuracy: What percentage of alerts are valid incidents vs. false positives? High false-positive rates desensitize teams and reduce alertness.

Response Metrics

Response time: How long between detection and first meaningful action? This measures organizational agility.

Escalation time: How long before the right people are engaged? Slow escalation extends outages unnecessarily.

Recovery Metrics

Mean time to recovery (MTTR): Now in proper context—not as the primary metric, but as one component of a larger picture.

Recovery success rate: What percentage of recovery attempts succeed on the first try? Rollbacks that fail twice are particularly damaging.

Impact Metrics

Customer impact ratio: What percentage of users actually experience impact from a given incident? A backend service down for 30 minutes might affect 0% of users if it's not on the critical path.

Revenue impact: How much revenue was at risk during the outage? This bridges technical metrics to business outcomes.

Learning Metrics

Postmortem completion rate: What percentage of incidents have postmortems completed? Target 100% for any incident affecting customers.

Action item completion rate: What percentage of postmortem action items are completed? Track this monthly. Consistently low completion rates indicate either unrealistic action items or organizational commitment problems.

Recurrence rate: What percentage of incidents are repeats of previous incidents? High rates indicate learning isn't happening effectively.

Building a Resilient Organization

Resilience isn't built through metrics alone. It requires systemic changes:

Invest in detection and observability: Spend on tools, training, and time to understand what's happening in your systems. Prevention is cheaper than recovery.

Build response muscle: Regular incident simulations and gamedays help teams execute under pressure. Like any skill, response capability improves with practice.

Invest in reliability engineering: Dedicate engineers to reducing system complexity, improving automation, and building resilience into architecture. This is not someone's part-time job.

Create a learning culture: Blameless postmortems only work if people believe they're actually blameless. Leaders must demonstrate this through action. Never punish learning. Instead, celebrate organizations that surface problems quickly.

Measure and monitor resilience: Use the metrics framework above. Review them monthly. Use them to guide investment and prioritization.

The Bottom Line

MTTR is a tactic; resilience is a strategy. While it's important to recover quickly from incidents, the real goal is to build organizations that prevent incidents, detect problems early, respond effectively, and learn continuously. When you focus on these broader dimensions, resilience becomes a competitive advantage that protects revenue, customer satisfaction, and team morale.

The most resilient organizations aren't those that fix things fastest—they're the ones that break things least often and learn most effectively.

About the Leadership Hub

The SRE Leadership Hub provides practical insights, frameworks, and advisory guidance for building reliable, scalable, and intelligent engineering organizations.

Related Articles

Reliability Is a Business Strategy

Explore how forward-thinking organizations are shifting reliability from an operational concern to a core business strategy.

Building Effective SLOs and Error Budgets

Learn how to define meaningful SLOs and use error budgets as a powerful tool for balancing reliability and velocity.

Get New Insights Delivered

Subscribe to our newsletter for engineering leadership perspectives and insights.