Retries turn an outage into an overload

When a region fails, its clients retry, and the retries land on whatever is still serving, on top of the traffic moved there. A retry buys one caller a better chance by spending more of the server's time, which is harmless when failures are rare and damaging under overload.

Layers multiply the effect. AWS's Builders' Library works through a five-deep call stack that retries three times at each layer: once the database at the bottom starts failing, its load rises 243-fold and it is unlikely to recover. A service coming back up, Sarah Wells notes, meets new requests and old retries at once. A survivor with headroom only for the moved traffic (Capacity must already be there) has none for that.

The defences limit how much retrying the whole system does: exponential backoff with a cap, jitter so retries stop arriving in synchronised spikes, a local token bucket that throttles retries once spent, and retrying at one point in the stack rather than at every layer. Failover automation is another reaction to errors worth the same scrutiny.

Retried writes carry a second risk, since a request that timed out may have succeeded: Idempotency keys make retries safe.