Resilient by Design: Building Systems That Expect Failure
Why technology resilience is only the foundation?and business resilience is the outcome.
Why technology resilience is only the foundation?and business resilience is the outcome.
Most engineering organizations are very good at designing systems that are highly available.
Multiple availability zones. Redundant databases. Automated failover. Backups. Disaster recovery environments. Observability. Incident response. On-call rotations.
All of these capabilities matter.
But they can create a dangerous illusion: that technology resilience and business resilience are the same thing.
They are not.
A system can be technically redundant and still leave customers unable to complete the transaction that matters to them. A platform can remain available while an essential business service is effectively unavailable. A recovery procedure can work exactly as designed while the business continues to absorb customer, financial, regulatory, or reputational harm.
Technology resilience is the means. Business resilience is the outcome.
Resilient by design means building systems that expect failure, contain it, absorb it, degrade intelligently, recover predictably, and learn from it.
Reliability engineering often starts with a reasonable question: “How do we keep the system available?”
Resilience requires a broader question: “What must continue to work when the system cannot?”
That distinction changes architecture decisions.
If the objective is infrastructure availability, teams naturally focus on servers, clusters, databases, networks and cloud regions.
If the objective is business resilience, the conversation moves upward: customer journeys, essential services, transaction flows, operational processes, people, third parties, recovery decisions and business impact.
The difference becomes visible during disruption.
When everything works, a highly available architecture and a resilient business can look almost identical.
When something fails, they separate quickly.
The real test of resilience is not what happens when everything works. It is what the customer can still accomplish when something doesn't.
Reliability and resilience are closely related, but they answer different questions.
Reliability asks whether a service performs consistently within defined expectations.
Resilience asks what happens when those expectations are violated.
Reliability tries to reduce failure probability.
Resilience assumes that some failures will still happen.
That leads to a useful leadership distinction:
Strong organizations need all three.
A resilient architecture does not attempt to eliminate every possible failure. It creates boundaries around failure.
The objective is not perfection.
The objective is controlled imperfection.
Technology resilience describes whether the technology estate can withstand disruption: infrastructure failures, database failures, network problems, capacity exhaustion, bad deployments, dependency failures, data corruption, cloud control-plane issues or regional disruption.
Business resilience asks whether the business can continue delivering its most important services despite those disruptions.
A useful way to see the relationship is:
Customer outcome
↓
Essential business service
↓
Customer journey
↓
Application / service
↓
Platform
↓
Infrastructure / dependencies / third parties
The customer does not consume infrastructure.
The customer consumes a service.
That means resilience has to be evaluated from the top of the stack down, not only from the bottom up.
Technology resilience is the foundation. Application resilience extends it. Service resilience makes it meaningful. Business resilience makes it matter. And customer resilience is the outcome we ultimately care about.
One of the biggest mistakes in resilience programs is starting with infrastructure inventories.
Teams catalogue applications, servers, databases, queues, APIs and cloud services. The inventory becomes extensive.
But the key question remains unanswered: Which business services must continue, and what level of disruption can the business actually tolerate?
Start with the service.
For a bank, examples might include access to money, payments, transfers, card transactions, trading, settlement, customer authentication and other critical financial services.
Not every function deserves the same resilience target.
A reporting dashboard being unavailable for two hours is fundamentally different from customers being unable to access their money on payday.
Resilience therefore requires prioritization.
For each essential service, leaders should understand:
This is where resilience becomes an engineering leadership problem rather than an infrastructure checklist.
The Barclays incident in early 2025 is a useful example because the technology failure and the business impact were tightly connected.
Barclays told the UK Parliament's Treasury Committee that the root cause of the January 31, 2025 incident was a software problem in a critical module of its UK mainframe operating system. The bank said the problem caused progressively severe degradation in mainframe processing performance.
That mainframe supported core banking applications including deposits, overdrafts and debit cards across Barclays UK, as well as services for other parts of the bank.
The incident therefore was not simply a “mainframe outage.”
It was a resilience problem at the boundary between a technology dependency and essential customer services.
The UK Treasury Committee subsequently reported that 56% of online payment attempts during the incident failed because of severe degradation in mainframe processing performance.
That distinction matters.
The important question is not: “Was the mainframe highly reliable?”
The more important question is: “What happens to essential customer services when a critical dependency becomes impaired?”
Barclays' broader evidence to Parliament also describes resilience mechanisms such as alternative channels, duplicated systems, stand-in processing, manual processing, identification of critical banking services, and senior accountability for those services.
The lesson for engineering leaders is not that legacy technology is inherently bad.
The lesson is that technology concentration becomes business risk when the failure of one component can cross multiple service boundaries.
Do not ask only where your single points of failure are. Ask which customer outcomes depend on them.
The October 2025 AWS disruption provides another important resilience lesson.
AWS reported that increased error rates in US-EAST-1 were triggered by DNS resolution issues involving regional DynamoDB endpoints.
Resolving that initial issue did not immediately restore everything.
AWS reported that a small subset of internal subsystems remained impaired. EC2 instance launches were subsequently throttled, and additional services experienced downstream effects.
AWS later described an impairment in the EC2 subsystem responsible for launching instances because of its dependency on DynamoDB. Network Load Balancer health checks were also affected, producing connectivity problems across multiple services.
This is the classic resilience trap:
Redundancy at one layer does not guarantee independence at the dependency layer.
Two services can be deployed independently and still share a dependency that creates a common failure domain.
This is why architecture reviews that only ask “Do we have redundancy?” are incomplete.
Leaders should also ask:
Resilience is not about having more components.
It is about having meaningful failure boundaries.
The July 2024 CrowdStrike incident demonstrates a different category of resilience risk: change itself can become a failure domain.
CrowdStrike's preliminary post-incident report described a problematic Rapid Response Content update that caused Windows systems to crash. The company identified a defect in its content validation process that allowed the problematic content to pass through.
The engineering lesson extends well beyond endpoint security.
A deployment pipeline can be highly automated, observable and fast—and still have an enormous blast radius if the change is allowed to reach too many systems before the failure becomes visible.
This is why resilient change management includes:
The deeper lesson is simple:
A deployment is not resilient merely because it can be rolled back. It is resilient when the blast radius is controlled before rollback becomes necessary.
Many organizations treat resilience as a recovery problem.
Something fails.
Engineers restore it.
Service returns.
Incident closes.
But recovery is only one stage.
A resilient architecture asks what happens before recovery is possible.
Can the service absorb the failure?
Can it operate in a degraded mode?
Can non-critical functions be sacrificed to preserve critical ones?
Can customers complete the most important part of their journey through another path?
Designing for failure means making those decisions before the incident.
Don't design only for the system you expect to have. Design for the business you need to protect when the system behaves unexpectedly.
Dependencies are often where resilience assumptions become weakest.
Teams know their direct dependencies.
They are often less certain about indirect dependencies.
A service may depend on an API that depends on a database that depends on a storage system that depends on an identity service that depends on a shared control plane.
Each individual dependency may appear reasonable.
The combined dependency graph may be fragile.
Resilience reviews should therefore examine dependency depth, common-mode failures, third-party concentration and recovery dependencies—not simply uptime commitments.
A vendor promising 99.99% availability does not make your business resilient if your service has no alternative path when that vendor fails.
Graceful degradation is frequently described as a technical capability.
In reality, it is a business decision expressed through technology.
When capacity is constrained, which functionality stays?
When a dependency is unavailable, which transactions are prioritized?
When the system is overloaded, which customers or operations receive protection?
These decisions should not be invented at 2 a.m. during an incident.
Leaders should establish them beforehand.
A resilient system may deliberately provide less functionality during disruption so that it can continue providing the functionality that matters most.
That is not failure.
That is resilience.
Financial services make this distinction especially visible.
Not every banking feature carries the same consequence when unavailable.
A customer being unable to change a profile preference is inconvenient.
A customer being unable to access money, make a payment, transfer funds, complete a settlement or execute a time-critical transaction can create immediate financial harm.
Trading is another example where timing can materially change the impact of disruption.
This means availability targets should not be assigned uniformly across the technology estate.
The resilience requirement should follow the business consequence.
Ask in this order:
What customer outcome matters?
What business service delivers it?
What level of disruption is tolerable?
What technology capability is required to protect it?
This is also why regulators increasingly frame operational resilience around important business services and impact tolerances rather than infrastructure availability alone.
One of the most dangerous statements during an incident is: “The system is back.”
Back does not necessarily mean recovered.
A database may be online while transactions remain queued.
An API may respond successfully while downstream reconciliation is incomplete.
A payment platform may be technically available while customers cannot complete their end-to-end journey.
Backlogs, stale data, duplicate transactions, manual work and customer support queues can continue long after infrastructure recovery.
Therefore:
Technology recovery ≠ Business recovery.
Recovery objectives should include business validation:
Technology is not the only failure domain.
People are part of the architecture.
If only one engineer understands the recovery sequence, that engineer is effectively a single point of failure.
If only one team can make a critical production change, the organization has created an operational dependency.
If recovery requires undocumented knowledge, manual commands and exceptional permissions, resilience is weaker than the architecture diagram suggests.
This is why resilient organizations invest in:
If recovery requires heroics, the system may be recoverable—but it isn't resilient.
Resilience is not free.
Redundancy costs money.
Multi-region architecture costs money.
Additional testing costs engineering capacity.
Isolation creates complexity.
Graceful degradation requires product decisions.
Manual fallbacks require operational readiness.
The leadership mistake is to ask: “How resilient can we make everything?”
The better question is: “Where does additional resilience create enough business value to justify its cost?”
This is a portfolio decision, not simply an engineering decision.
A practical way to make this trade-off visible is to establish a resilience budget for each essential business service.
| Dimension | Leadership question |
|---|---|
| Availability | How much disruption can the service tolerate? |
| Degradation | What functionality can be sacrificed? |
| Recovery | How quickly must the business outcome return? |
| Dependencies | Which common failure domains can affect the service? |
| People | Can recovery succeed without specific individuals? |
| Third parties | What happens if a critical supplier fails? |
| Data | What data loss or inconsistency is acceptable? |
| Customer impact | What harm occurs while recovery is underway? |
This creates a more useful conversation than simply asking whether a system is “highly available.”
Resilience that has never been tested is an assumption.
Teams often test individual components:
Those tests are necessary.
They are not sufficient.
The stronger test is end-to-end: Can the essential business service still deliver its intended customer outcome under a severe but plausible disruption?
That might mean testing:
The test should follow the customer journey all the way through.
If the customer outcome fails, the business service is not resilient enough, regardless of how healthy the underlying infrastructure metrics look.
A simple five-stage model can help engineering and business leaders evaluate resilience consistently:
Prevent
Reduce the likelihood and severity of failure through engineering
controls, testing, safe change and capacity planning.
Absorb
Keep local failures from becoming systemic failures through isolation,
redundancy, dependency boundaries and controlled blast radius.
Degrade
Preserve the most important customer and business outcomes when full
functionality is impossible.
Recover
Restore service and business capability safely, predictably and with
controlled recovery procedures.
Learn
Convert failures and exercises into architectural, operational and
organizational improvements.
Most organizations are strongest at Prevent and Recover.
The biggest resilience opportunities often exist in the middle: Absorb and Degrade.
That is where architecture meets business prioritization.
Leaders should be able to answer a small number of uncomfortable questions without needing a slide deck.
The last question is particularly important.
A resilience exercise that produces no architectural, operational or organizational change may be an exercise in compliance rather than resilience.
Resilience becomes powerful when it stops being owned exclusively by the infrastructure or SRE organization.
Product leaders need to define which customer outcomes matter most.
Engineering leaders need to design the technical paths that protect them.
Platform teams need to provide resilient building blocks and reduce common failure modes.
SRE teams need to provide reliability expertise, standards, testing, operational insight and leverage.
Business leaders need to make explicit trade-offs around impact, cost and acceptable disruption.
No single team can create business resilience alone.
But someone must own the outcome.
That is the leadership shift:
Move the resilience conversation from “How do we keep our systems running?” to “How do we keep the business delivering what matters?”
Every organization eventually experiences a failure it did not expect.
The question is not whether the architecture contains every possible failure mode.
It cannot.
The question is whether the organization has designed enough resilience around the failures that matter most.
When the next dependency fails, can the business absorb it?
When the next deployment goes wrong, can the blast radius be contained?
When the next critical platform becomes unavailable, can customers still complete the most important part of their journey?
When recovery takes longer than expected, does the organization know what to protect first?
And when the recovery requires action, can the team execute without relying on heroics?
Those are resilience questions.
The leadership question is simple:
If a critical technology dependency failed tomorrow, what essential business service would fail with it—and what have we deliberately designed to keep working?
Resilience is not the absence of failure.
It is the ability to continue delivering what matters when failure occurs.
Technology resilience provides the foundation.
Business resilience determines whether that foundation protects customers, revenue, trust and critical business outcomes.
The strongest engineering organizations do not simply build systems that survive failure.
They build organizations that know what must survive.
Build for failure. Design for degradation. Test the business service. Protect the customer outcome.
That is what resilient by design really means.
Explore more practical perspectives on reliability, engineering leadership, platform strategy and operational excellence.
Explore Insights