June 11: ORD transit loss destabilized Managed Postgres clusters
#June 11: ORD transit loss destabilized Managed Postgres clusters (01:18UTC)
We hit a period of severe packet loss and intermittent connectivity on network transit paths into/out of ORD, which made some Managed Postgres nodes unable to reliably reach Patroni’s DCS (Kubernetes). That triggered leadership churn (primaries demoting / failovers) and left some replicas unable to participate cleanly, causing intermittent connection errors for Managed Postgres clusters in ORD. Service stabilized once reachability improved, and we also rescheduled affected replicas away from the worst-impacted hosts to keep clusters steady during ongoing network flaps.