This design tolerates a regional outage
Every word in this sentence is easy to understand, and it still leaves open how the design copes, under which conditions, and what happens to writes while a region is gone. It reads as one claim. It is at least four, and they can fail independently.
- The service keeps answering when a region goes away: Availability and durability are different promises.
- It keeps the data it had already accepted: An acknowledged write is a promise to the user.
- It recovers within a time someone agreed to: RTO and RPO name the two costs of an outage.
- Something decides the region is gone, and decides correctly: Failover is a decision, not an event.
Beneath all four sits an assumption that is easy to leave unstated: that a usable copy exists in another region. On AWS, data created in one region exists in no other unless something replicates it. How current that copy is, and whether anything can safely promote it, is where most trails through these notes end up. A good place to start pulling is A replica is only as current as its lag.
Two questions come before any of the detail. Is a regional outage even the failure this design is most likely to meet? Most outages are not regional suggests often not. And has anyone watched it survive one? Until someone has, the claim is a statement of intent.
Follow a link and its note opens beside this one. Keep going and earlier notes fold to their titles at the left edge; click a title to go back.
These notes were written by a language model, drawing on published books, documentation and incident reports, and have not been reviewed by a person.