September 23: Our gateways took out 6PN for a bit
#September 23: Our gateways took out 6PN for a bit (14:44UTC)
We broke private networking (6PN) for a certain subset of machines. The root cause was a collision between how we assign internal IP addresses to gateway nodes and how we route 6PN traffic to machines.
To recap a bit how 6PN works: when a machine is first created, it is assigned a 6PN address in the form fdaa:netid:netid:hostid:hostid:machid:machid:2. The hostid:hostid octets are derived based on the physical host where the machine is located at the time of creation (or on migration, before ~May 2025). It is also important for routing purposes: even after we made 6PN addresses stable back in 2025, for machines that have never been migrated, the custom eBPF program we use for 6PN routing still by default makes routing decisions solely based on hostid.
This year, we have been working on improvements to our 6PN setup, and part of this plan is to remove more of our custom routing implementation, which has been hard to debug and make any changes to at all, and lean on standard Linux routing tables. This also gives us much higher flexibility in defining how packets are routed internally in our network, unblocking a lot of reliability and performance work, which you should hear about in separate updates soon. An important part of this work is to decouple host IDs entirely from 6PN routing: machines’ routes are now determined entirely by where they are currently located, not how their addresses happen to be assigned.
One headache to deal with in this new reality is gateways. There are two types of gateway peers: static and interactive. Static peers are those you would get from fly wireguard create, which you can import into a standard Wireguard client to connect into your organization’s 6PN network. Interactive peers are those used by flyctl for transient operations such as fly ssh console. The difference is that static peers have a persistent identity in our platform’s routing state, while interactive peers do not. They can only be routed based on looking at the host ID, which, in this case, will be the gateway’s host ID.
There are no other flags in the address for us to distinguish gateway peers from machines, so the easiest solution here was to just route the entire gateway’s /80 prefix (per netid) to the gateway host by adding such a rule in Linux’s routing table. This worked perfectly fine for our current gateways! However, one separate line of work was to make our edges act as gateways for non-static peers. This means that every edge had to be treated like a gateway in our 6PN routing logic, and this was the direct cause of this incident:
When a physical worker host is decommissioned, its internal IP address is released. This also determines its host ID. A newly provisioned host can either pick up a brand new IP / host ID or one from a decommissioned host. Ever since 6PN addresses became stable, this fact means that it is possible for machines to retain host IDs that end up being re-assigned to a non-worker host. This is exactly what happened when we started treating edges as gateways: they get /80 routes, which ended up overriding the code that handled per-machine routing (it depends on machines that have not yet been accessed locally not having a route). Remember, machines that have those edges’ host IDs (by retaining 6PN from the original worker host) are most certainly NOT located on the edge (they have to be on workers by definition!).
We mitigated the first occurrence within about 10 minutes by deleting the bad /80 routes from the routing table on all affected hosts and rolling back changes to s6pnforwarderd, the daemon that manages 6PN routes. We then shipped a fix by only ever adding interactive peers as /128. This, unfortunately, did not fix the root cause of the incident: it is not just the /80 routes that are the problem but the priority between the gateway logic and worker logic. This caused a brief recurrence of this issue the following day. We then fixed the root cause at least for the moment by always looking for any machines that may bear a host’s ID first, regardless of whether the host is known to be a worker or not.
This incident exposed a lot of issues with our 6PN setup. First and foremost, not everybody even realized as a first reaction that a worker’s host ID may be reused by a non-worker host. This is not surprising if one knows how host ID assignment worked internally, but is really not well-documented. Our gateways have not been touched for a long time, and this is the only reason the issue did not surface until we started using edges as gateways. We have provisioned a lot of new edges recently to grow our capacity and for better connectivity, and that is why they ended up reusing decommissioned workers’ host IDs.
A deeper issue is that we need better alerting and automated testing for 6PN routing. All of our existing checks use machines sitting on their original host, which means they’d never catch a conflict caused by host ID reuse, or really any issue at all related to the stable 6PN address implementation. It is, in our opinion, quite amazing that we only got bit by this right now, which is definitely not a good thing. Many, many problems could have sneaked under our radar because they didn’t cause an alert or a more widespread outage like this. As part of the 6PN reliability improvement project, our plan is to, for now, at least assign non-local 6PN addresses to machines we use to test and alert on 6PN issues. The medium-to-long term goal is that host IDs in 6PN addresses should carry exactly no weight in deciding how they’re routed. This is already (sort of) true today for machines, but not true for gateway peers. It needs to be true for 6PN in general, lest we leave the door open for this type of incident to recur. The current state is a hack to keep existing gateway peers working (since we can’t force all existing peers and machines to a new 6PN address), but we really need to transition to a better model across the board for new peers and machines.