beginner 2 min answer Multiple choice

At 09:00 a team creates the DNS record for a new regional endpoint. It resolves from their laptops at once, but part of the fleet keeps getting NXDOMAIN until about 09:20 even though the record's TTL is 60 seconds. Why?

dnsnegative-cachingnxdomaincutoversoa
Pick one
Show the full answer Hide the answer

The mechanism

Resolvers cache failures as well as successes. RFC 2308 (1998) defines negative caching: when an authoritative server answers "this name does not exist", the resolver stores that answer for an interval taken from the zone's SOA record - in practice the lesser of the SOA MINIMUM field and the SOA record's own TTL. Managed zones commonly ship a default in the 300 to 900 second range, which is where the 20 minutes comes from.

The record's own 60-second TTL governs positive answers, and only once a positive answer exists to cache. Any resolver that was asked for the name before it existed is serving a cached no, and it will keep doing so until that interval expires. Which resolvers those are depends on who happened to ask early: the laptops asked after creation and a health check or a config reload asked before.

Why the other options fail

  • TTL starting after the first lookup. A TTL is a lifetime attached to a record in a response. It is not a clock that starts somewhere globally, and this model also predicts the wrong fix.
  • A fifteen-minute resolver minimum. Resolvers do clamp TTLs they consider too low, but there is no fifteen-minute floor, and believing in one leads to raising the TTL - which changes nothing here.
  • Propagation between authoritative servers. Zone transfer to secondaries is usually seconds, and many managed services do not work that way at all. It also would not produce a consistent NXDOMAIN per resolver for twenty minutes.

What to do before a cutover

  1. Create names early, pointing at the current target or a placeholder, so no resolver ever caches a negative answer for a name the cutover depends on. This costs nothing and removes the whole class of problem.
  2. If you cannot pre-create, lower the SOA minimum a day ahead and restore it after.
  3. Never put record creation inside the cutover window. The symptom - half the fleet working, half returning NXDOMAIN - reads exactly like a partial outage and sends the response in the wrong direction.

When this is the wrong thing to change

Negative caching exists for a reason: it absorbs floods of queries for names that do not exist, which is most of the junk traffic an authoritative server sees. Do not set it near zero to make cutovers convenient. For a name that already exists and is only changing its target, negative caching is irrelevant and the TTL is the number to argue about.