Skip to main content

Resources & Guides

Practical frameworks, playbooks, checklists, and reference architectures to support your engineering transformation.

SRE Frameworks

Featured Framework

SRE Leadership Framework

A practical framework for engineering leaders to build reliable, resilient, scalable, and cost-efficient engineering organizations at scale.

Explore the Framework →
Leadership Operating System

SRE Leadership Operating System

A practical leadership system connecting reliability strategy, SLOs, resilience, platform engineering, observability, AI-assisted operations, engineering economics, and organizational learning.

Explore the Operating System →
Framework

SRE Operating Model Framework

Comprehensive framework for establishing SRE practices, defining roles and responsibilities, and building reliability into your engineering organization.

Framework

Incident Management Framework

End-to-end framework for incident detection, response, investigation, and learning including post-mortem processes.

Reliability Checklists

Checklist

Production Readiness Checklist

Essential checklist items to ensure services are production-ready before deployment. Covers architecture, operations, security, and monitoring.

Checklist

Disaster Recovery Readiness

Checklist for ensuring your disaster recovery plans are comprehensive, tested, and ready for execution.

Engineering Playbooks

Playbook

Incident Response Playbook

Step-by-step guide for incident detection, response, investigation, and post-mortems with templates and best practices.

Playbook

On-Call Runbook

Comprehensive guide for managing on-call rotations, escalation procedures, and incident triage.

Leadership Guides

Guide

Key Engineering Metrics Guide

Framework for selecting, measuring, and acting on engineering metrics that matter for business and operational outcomes.

Guide

Engineering Operating Model Design

Guide for defining organizational structures, processes, and governance that support engineering excellence and reliability.

Cloud Architecture

Reference

Cloud Architecture Reference Patterns

Proven architectural patterns for building scalable, reliable, and secure cloud systems with best practices.

Reference

Multi-Cloud Strategy Framework

Framework for evaluating, selecting, and managing multi-cloud strategies and architectures.

Incident Management

Guide

Disaster Recovery Planning

Comprehensive guide to disaster recovery strategy, planning, testing, and organizational readiness.

Guide

Root Cause Analysis Guide

Best practices and techniques for conducting effective root cause analysis and creating actionable remediation items.

Observability

Guide

Observability Best Practices

Learn how to instrument applications and infrastructure for comprehensive observability, monitoring, and alerting.

Framework

SLI/SLO Definition Framework

Framework for defining Service Level Indicators and Objectives with practical examples and case studies.

FinOps

Guide

FinOps Cost Optimization Strategy

Framework for implementing FinOps practices and optimizing cloud spending across your organization.

Checklist

Cloud Cost Optimization Checklist

Quick checklist for identifying and implementing immediate cost savings in cloud infrastructure.

Need Custom Guidance?

Our advisory team can help develop frameworks and resources tailored to your organization's unique needs.

Start a Conversation