Automatic failover can make an outage worse
Automation shortens the time to act, and a brief network interruption is exactly the event it is most likely to misread.
GitHub's October 2018 incident is a clear public account. Replacing failing optical equipment broke connectivity between their US East Coast network hub and primary East Coast data centre for 43 seconds. In that window, Orchestrator nodes elsewhere formed a quorum and began moving database primaries to the West Coast. The East Coast servers held a short run of writes that had never replicated, and for more than half an hour afterwards the West Coast took writes of its own. Service stayed degraded for 24 hours and 11 minutes.
The tooling had treated a partition as a failure, the situation described in A partition looks like an outage from inside. Promotion across regions with asynchronous replication stranded writes on the old side (Promoting a replica can lose acknowledged writes), and two histories then had to be reconciled. GitHub notes that Orchestrator behaved as configured; the application tier could not cope with cross-country write latency.
AWS's disaster recovery guidance urges caution with failover triggered by health checks. GitHub's own fix was to stop Orchestrator promoting primaries across regional boundaries: automation that knows what it may not do alone.