Staged rollouts contain a bad change

Regions that take the same change at the same moment can fail together. Releasing to a small share first, then widening only while the metrics hold, keeps a bad change to whoever received it and turns a global failure into a partial one (Most outages are not regional). AWS does this internally: deployments to the zones of a region are separated in time so one update cannot take them all down.

The weak spot is whatever bypasses the rollout. Cloudflare's software passed through an employee-only site, a slice of free traffic and a few canary sites before going global, but its procedure let firewall rules go everywhere at once, for speed against new threats. That exception caused the July 2019 outage, and afterwards rules were staged like other software, with an emergency global path kept for active attacks. Sarah Wells argues that configuration changes especially should be staged. A data change travels by replication, which stages nothing.

A canary that alters a shared resource can do damage even at 1% of traffic, because the resource is not staged with it. Failing over is a change too, and deserves the same care (Test the failover or you have not got one).