August 14: Petsem primary crash loop from corrupted volume

August 14: Petsem primary crash loop from corrupted volume (19:26UTC)

The primary node of Petsem (our secrets database) suffered from disk corruption during a routine deployment. Petsem can run in a degraded state when the primary is unavailable, so while creating new apps and updating secrets failed during the impact period, it was still possible to perform read actions, such as deploying new machines.

We restored service by promoting an up-to-date replica to become the new primary, repointing the other replicas to follow it, and then provisioning a new replica to return the cluster to full redundancy.