▪▪▪ ▪▪▪

What To Actually Read About Anycast

by Robert Jefe Lindstaedt September 19, 2026
Essential Reads

Anycast takes one sentence to describe: many sites announce the same address, and each client reaches the nearest one. Everything hard is in what happens when a site leaves.


I went looking for that second part this week, working out why taking one site out of rotation was visible to users at all. Every link below was fetched and every quote read in the source, not in a summary.

Start here

RFC 4786 / BCP 126, Operation of Anycast Services is the closest thing to a manual. It is a best current practice, not a protocol, and it spends its length on production problems: how to take a node out (§4.8.2 covers pessimistic withdrawal, where a covering prefix is withdrawn “if any individual service becomes unavailable”), why you want a hold-down before re-advertising, and how to monitor. That last part is the one worth pinning up, and the closing section below is about why.

RFC 7094, Architectural Considerations of IP Anycast is the conceptual companion: what anycast is, and what it does to stateful protocols. §4.5 restates RFC 3258’s advice that routes “should not be withdrawn, in order to reduce operational complexity”, then argues it deserves more attention.

The older, narrower ones

RFC 1546, Host Anycasting Service is the original definition, and worth ten minutes because it is honest that anycast is a stateless idea being used for stateful things.

RFC 3258, on distributing authoritative name servers, is DNS-specific but holds the most useful drain advice in the set (§2.3): do not withdraw the route, shut the service process down instead.

When you tear a BGP session down

If your drain works by dropping a BGP session, read RFC 4724 (graceful restart) and RFC 8538 (the same for sessions that end in a BGP notification message) before you trust that the withdrawal happened. Graceful restart keeps a neighbour’s routes alive across a restart (RFC 4724 §4.1), so the neighbour may keep sending traffic to a box you just took away. RFC 8538 sets a default stale timer of 180 seconds and lists administrative shutdown as a hard reset that should flush at once. Whether it does is a property of two implementations, not of the RFC.

Running it at scale

The vendor write-ups state the tradeoffs plainly. Dropbox on their edge network has the sentence everyone eventually rediscovers: “graceful drain of the PoP is impossible”, because “BGP balances packets and not connections”. For what a real graceful drain costs to build, Fastly’s faild describes the flow-tracking layer it takes. Cloudflare’s Traffic Manager argues for not reaching for withdrawal at all, and their anycast primer is the plain-English version for anyone you need to explain this to. The root server operators publish their practice: K-root and F-root.

Easy to state, hard to verify

The trap is measurement, not design. A global routing view looks like proof that nothing happened. It is not: if a covering prefix stays announced from your other sites, a collector sees no change while one region is dark, because the withdrawal never left the provider’s network.

The instruments nearest to hand each answer a different question. A collector reports the global table, a probe reports its own catchment, an uptime check reports whichever site it reached. None watches the region you are changing, which is the only place the change can hurt.

RFC 4786 §5.1 has asked for probes “distributed representatively across the routing system” since December 2006, with the node’s identity recorded alongside each answer. Cloudflare does the second part in public: curl https://www.cloudflare.com/cdn-cgi/trace names the site that served you, and they published how far you get inferring catchments with no probes.