Replication lag peaks during incidents
The lag figure people quote is measured on ordinary days. The one that matters is the lag at the moment the primary's region fails, and that moment is rarely ordinary.
The conditions that stretch replication are the conditions of an incident. Transient network faults and lost packets lengthen the time a change takes to reach every copy, and heavy load on a replica is probably the largest cause of delay. A region in trouble is likely to show both, just as someone is trying to decide whether to fail over.
GitHub's 2018 outage shows how far this can go. A 43-second network break led to a cross-country failover and 24 hours of degraded service. During recovery, dozens of read replicas were hours behind their primary, and as users in Europe and the US began their working day the delays grew instead of shrinking, so many requests read data hours old.
An average hides this. The useful figures are the tail over a long period and the lag during past incidents, which exist only if someone kept them. The honest way to learn the real number is to fail over on purpose, under load, and count what was missing: Test the failover or you have not got one.