Failure detectors can only suspect

Every automated failover begins with something concluding that a region has failed, and that conclusion is an inference from silence. In practice it is a timeout: no heartbeat or health check reply for long enough.

A caller that hears nothing cannot tell a crashed server from a slow one or from a lost reply; from its side these look identical. MongoDB, for instance, starts electing a new primary when a secondary has missed heartbeats for ten seconds by default. Set a timeout like that short and ordinary blips count as failures, each one paying the cost of promotion; AWS warns that failing over on a false alarm still costs downtime and data. Set it long and every real failure waits that much longer, straight out of the RTO.

No single setting serves both. The timeout works better as the start of a decision than as the decision itself: suspect quickly, then confirm from somewhere else. AWS recommends canaries running from the standby Region, since the failing Region's own monitoring may be impaired too. Only then does the decision begin.

A detector trusted to act on its own suspicion will sometimes act on a partition, which is one way Automatic failover can make an outage worse.