June 10: Two GRU edges OOM

June 10: Two GRU edges OOM (19:35UTC)

Two edge nodes in the GRU region couldn’t keep up with exporting their metrics, and the backlog caused the hosts to run out of memory. The immediate impact lasted about 17 minutes, at which point we rebooted the affected nodes and returned them to the routing pool. Some traffic was disrupted, though this was not a complete outage in the region. We’ve tidied up our metrics pipeline a little since, paring back some high-cardinality metrics that contributed to the heavy load.