Resilience tests should include control-plane failure
Most resilience testing assumes you can still manage the platform. Sometimes you cannot.
The provider's management layer being degraded means you cannot scale, deploy, fail over or sometimes even see what is happening, while your workloads continue running or not. It has happened. Testing only application failure leaves the scenario where your tooling is what is broken entirely unrehearsed.
More on Cloud resilience
- Multi-region is not the same as multi-providerTwo towns, one mains
- Autoscaling can scale a failure tooIt copied the mistake
- Managed services move responsibility, not all riskThe van takes the machines
- Private networking does not automatically mean trustedNobody asks twice
- Egress is part of the cloud security boundaryWatch the loading bay
- Snapshots can copy sensitive data silentlyThe copy keeps nothing
