practice

Route Origin Validation

also called ROV, RPKI Origin Validation

Checking a BGP announcement's origin against a signed authorisation before accepting it, so a more specific prefix announced by the wrong network is dropped instead of winning.

bgprpkiroahijackrouting-securitymaxlength

BGP accepts what it is told. If another network announces a /24 inside your /16, every network that hears it sends your traffic there, because forwarding follows the longest matching prefix and takes no interest in who owns the address space or how long the path is. The protocol has no authorisation to appeal to.

In 2008 a national telecom tried to block a large video site inside its own borders by announcing a more specific prefix. The announcement leaked to its transit provider, propagated globally, and took the site off much of the internet for roughly 2 hours. The intent was a local filter and the mechanics were those of a hijack, which is the usual shape: operator error with the blast radius of an attack.

RPKI (RFC 6480, 2012) supplies the missing authorisation. The address holder publishes a signed Route Origin Authorisation naming which autonomous system may originate a prefix and up to what length. Routers fetch validated payloads from a local validator and mark each route Valid, Invalid or NotFound. Validation alone changes nothing; dropping Invalid routes is what turns the data into protection.

Why it matters

A hijacked prefix is not only an outage, it is a credential: attracting your traffic is enough to pass domain-validated certificate issuance and then present a trusted certificate for your name, which is how hijacks have been turned from disruption into theft.

Deployment has also inverted the risk. About 65% of announced IPv4 prefixes carried a ROA by 2026 according to public RPKI monitors, and most large transit networks and exchange route servers drop Invalids. Your likeliest routing outage is therefore no longer somebody hijacking your prefix; it is your own announcement being Invalid because the ROA was not updated.

Implementation patterns

  • One ROA per aggregate you actually announce, with the maximum length set to the longest prefix you genuinely originate. A blanket maximum length of /24 over a /16 authorises any /24 inside it, which is an invitation rather than a control.
  • Create the ROA before the announcement changes, as part of the turn-up checklist for a new site or transit provider.
  • Run at least two validators feeding routers over RPKI-to-Router, configured so absent data means NotFound and NotFound is accepted, so a dead validator fails open rather than withdrawing the internet.
  • Monitor your own prefixes from outside, because only a looking glass or a validation monitor tells you whether the world considers your announcement Valid, and stage enforcement: log what would be dropped for several weeks, review the list, then enforce.
  • Remember the gap. Origin validation says nothing about the AS path, so a hijack announcing the correct origin behind a forged path still succeeds; that needs ASPA or BGPsec.

Industry example

The grounding is industry-wide rather than one company's blog. The 2008 incident is documented in detail by the operator community, the MANRS programme (2014) codified ROA publication and Invalid filtering as baseline expectations for networks and exchanges, and public monitors have tracked coverage climbing past half of announced IPv4 prefixes towards two thirds by 2026. The consistent lesson is that the recurring incidents are mis-set maximum lengths and missing ROAs rather than cryptographic failures.

Failure scenarios

  • Maximum length set too loosely, which signs away the very attack being defended against, since any more specific prefix inside the range validates.
  • A new announcement with no ROA, dropped by the validating networks you most wanted to reach while a traceroute from a non-validating provider looks perfect. This is the most common self-inflicted RPKI outage.
  • An expired or deleted ROA after a key rollover or registry account change, turning a working prefix Invalid globally with no change on your side.
  • A router configured to drop NotFound rather than Invalid, which discards most of the internet the moment a validator is unreachable, or a stale validator serving pre-revocation data, which accepts a hijack for as long as the cache lives.

Trade-offs

Choose Gains Pays
Publish ROAs Your prefixes become hard to hijack wherever filtering exists A dependency on registry CA infrastructure and a step before every routing change
Drop Invalids Hijacks and leaks with a wrong origin stop at your border Validator infrastructure, and the risk of dropping a legitimate route whose owner mis-signed it
Publish but do not validate Nothing to operate; others protect your traffic You still accept hijacks of everyone else's space

When not to use it

A stub network with one upstream and a default route gains little from validating, because it has no alternative routes to choose between, though it should still sign its own space. Do not enable Invalid-drop without the logging period first, since the first thing it drops is usually your own mistake. And do not use maximum length to pre-authorise prefixes you might announce one day: that is the single setting whose looseness cancels the mechanism.

Decision rule: sign everything you announce with a tight maximum length, validate at every border where you have a choice of routes, and gate enforcement behind a logging period with a named owner for the would-be-dropped list.

Interview question

Q: Your company announces a /16 from one data centre and is turning up a second site announcing a /20 from a different autonomous system number. After turn-up the new site is unreachable from several large networks. What happened, and what is your checklist for the next site?

What a strong answer covers: that the /20 is Invalid because the existing ROA authorises only the first AS and a maximum length that excludes it, so validating networks drop it while others accept it, which explains partial reachability and a clean traceroute from your own transit; that the fix is a ROA covering the /20 and the new origin AS, published and propagated before the announcement; that verification belongs outside the estate, from a looking glass and a validation monitor; and that the same checklist covers ROA expiry and registry account ownership, because those produce the identical symptom months later.

Quick check

Quiz: Why does a more specific prefix announced by another network capture your traffic even when your path is shorter? Longest-prefix-match is decided before any path comparison, so a /24 beats a covering /16 wherever it propagates.

Flashcard: What is the most common RPKI-caused outage, and which ROA field does the most damage when set carelessly? — Your own new announcement being Invalid because the ROA was not updated first; the maximum length field, which when loose authorises any more specific prefix inside your range.