Resilience policy  / field guide
Practitioner field guide · 29 September 2026

The decade Netflix took resilience out of the client library

Between 2016 and 2026 almost every Netflix library that decided something about the network was retired, while the libraries that decide something about data survived. This guide reconstructs that decade from the repositories themselves, and turns it into a rule for deciding where a policy belongs in a system you are designing now.

22 primary sources 3 organisations' repositories 4 incident reports Evidence through September 2026 Reference document, not a sitting
01

The territory

A company changed its mind about where a rule should live. The record of it is in thirty repositories, and it is more specific than any retrospective.

Here is the problem without a product name attached to it. A large fleet of processes has to obey rules about the network: which instance to call, how many calls may be in flight, when to stop calling, what to do when the answer is slow. Those rules have to change faster than the processes do, because traffic shifts, capacity changes and dependencies fail on their own schedule. If a rule lives inside a library that every process links, then changing the rule means redeploying the fleet, and the number of teams who must agree to that is the real cost of the design. The question this guide answers is what happens to an architecture over ten years when that cost is paid repeatedly, and what the organisation does about it once it notices.

Netflix is the case worth excavating because its 2016 answer was the one the industry copied. Spring Cloud Netflix, which is how most organisations actually adopted it, advertised eight capabilities on its 1.4.x branch: discovery, an embedded discovery server, a circuit breaker and its dashboard, a declarative REST client, a client-side load balancer, a configuration bridge and a routing filter. Seven of the eight are rules about the network, and every one shipped as a Java library linked into the application process. On the current branch the same file lists two, and both are Eureka.

8 → 2
Capabilities advertised by Spring Cloud Netflix, 1.4.x branch against the current one
2021
Last commit to touch Hystrix core source, four years before the repository's newest commit
4 days
Between a discovery client's refresh loop dying and anyone noticing, in the one detailed incident report in this corpus
5 years
That the successor library's central design question has sat unanswered in its issue tracker
What this guide is built from, and what it is missing

The session that researched this page reached three hosts: github.com, raw.githubusercontent.com and gist.github.com. Every other host was refused by a network egress policy, including Netflix's own engineering blog, the conference and paper archives, and the trade press. So there is no blog post, talk, paper or cost figure in this corpus, and Netflix's own narrative of these changes is absent. What remains is the layer that is hardest to spin and easiest to date: deprecation notices, archive banners, commit histories, release trains and issue threads. The guide is written entirely from those, and every place where they run out is marked.

The finding that changed how this guide is organised is not that the libraries were retired. It is which ones. Netflix in September 2026 still ships Spectator, Hollow, the DGS GraphQL framework, Metaflow, Atlas, EVCache and Zuul, all updated within the last week of the organisation's repository listing. Hystrix, Ribbon, Archaius, Servo, Governator, the Simian Army, Conductor, Falcor and Titus are absent from the first thirty entries. Sort by what each thing decides and the split is clean: the survivors encode a data or instrumentation contract, the casualties encoded a rule about the network. Sections 2 and 3 test that against the stated reasons.

Figure 1 · What each Netflix component decided, and whether it survived

all retired by 2026

all maintained in 2026

Data and telemetry

Spectator
metric format

Hollow
dataset transport

DGS
GraphQL schema

Network decisions

Hystrix
circuit breaking

Ribbon
load balancing

Archaius 1.x
runtime config

Policy moved out:
gRPC interceptors,
adaptive limiters,
Zuul at the edge

Contract stayed put:
still a library
in the application

all retired by 2026

all maintained in 2026

Data and telemetry

Spectator
metric format

Hollow
dataset transport

DGS
GraphQL schema

Network decisions

Hystrix
circuit breaking

Ribbon
load balancing

Archaius 1.x
runtime config

Policy moved out:
gRPC interceptors,
adaptive limiters,
Zuul at the edge

Contract stayed put:
still a library
in the application

Every component in the red group decided something about the network from inside the application process, and none is maintained in 2026; every component in the green group encodes a format or a contract, and all of them are. Compiled from the Netflix organisation repository listing and each project's own status notice.
Diagram source

One related excavation already exists in this collection and asks a different question of the same kind of evidence: The half-life of in-house infrastructure reads a decade of Uber's repositories to classify how in-house systems exit, as donated, deprecated or drifted. This guide takes the exits as given and asks what replaced the thing, and what the replacements have in common. Where the two overlap, on how to read an archive banner and what a commit date is worth, section 8 corrects the earlier method rather than repeating it.

02

How it is actually built

The 2016 shape, the 2026 shape, and the three places the policy went.

The 2016 architecture has an unusual property: there is no control plane. Discovery, load balancing, circuit breaking and configuration all live in the calling process, as objects configured by properties that ship with the application. The registry server and the dashboard exist, but they are data sources for in-process logic rather than places where a decision is made. That is why the design scaled so well at first and became expensive later. Every call is decided locally, which needs no coordination; every change to how calls are decided needs every caller to redeploy.

The 2026 shape, reconstructed from what Netflix retired and what it points at instead, puts the same four decisions in three different places. Load balancing and its associated cross-cutting concerns moved into RPC interceptors: the Ribbon README states the destination directly, that the team "started building an RPC solution on top of gRPC" for "multi-language support and better extensibility/composability through request interceptors". Circuit breaking moved into a control loop that sets its own threshold, described in the next paragraph. Routing and edge policy stayed in Zuul, which is not a library at all but a gateway you deploy, and which is one of the few 2016 components still being committed to in September 2026. Configuration, in the Archaius sense of a library that reads properties into typed objects, largely disappeared as a separate concern, because a control loop that infers its own limits does not need a property to read.

Worth being precise about what did and did not move, because the popular version of this story is wrong. Netflix did not move resilience out of the process into a sidecar. An interceptor runs in the same process as the caller, in the same language, on the same thread. What changed is not the location of the code but the location of the decision: an interceptor is shipped and versioned with the RPC stack rather than written per-application, and a control loop takes its setting from measurement rather than from a property file that somebody has to be right about. The code stayed in the application. The judgement left it.

Figure 2 · Where each decision lives, 2016 against 2026

2026: three destinations

2016: in the application

no setting left
to configure

same process,
different owner

Discovery client

Load balancer

Circuit breaker
tuned thresholds

Config bridge

RPC interceptor
shipped with the stack

Adaptive limiter
infers its own threshold

Zuul gateway
deployed as a service

2026: three destinations

2016: in the application

no setting left
to configure

same process,
different owner

Discovery client

Load balancer

Circuit breaker
tuned thresholds

Config bridge

RPC interceptor
shipped with the stack

Adaptive limiter
infers its own threshold

Zuul gateway
deployed as a service

Notice that the 2026 column has fewer settings, not fewer components: the discovery registry survives, the circuit breaker is replaced by something that computes its own threshold. Reconstructed from the Ribbon status notice and the concurrency-limits README.
Diagram source

The replacement for the circuit breaker deserves reading in the original, because its README contains the clearest statement of the problem anyone involved wrote down. Netflix's concurrency-limits project argues that operators think in requests per second, set a limit below a measured tipping point, and then discover that "in large distributed systems that auto-scale this value quickly goes out of date and the service falls over by becoming non-responsive as it is unable to gracefully shed excess load". Its proposed alternative is to reason in concurrent requests instead, via Little's Law, and then to refuse to configure even that: the limit is estimated per node by treating it as a TCP congestion window and running a delay-based algorithm over observed latency. The Vegas variant estimates the queue as L * (1 - minRTT/sampleRtt) and moves the limit by one each sampling window. The design concedes its own motivation in a single line: "For large and complex distributed systems it's impossible to know all the hard resources."

Figure 3 · The control loop that replaced the configured threshold

yes

no

latency sample

queue < alpha

queue > beta

Inbound
requests

In flight
< limit?

Handler

Reject:
UNAVAILABLE

Estimate queue
L x 1 - minRTT/sampleRtt

limit + 1

limit - 1

yes

no

latency sample

queue < alpha

queue > beta

Inbound
requests

In flight
< limit?

Handler

Reject:
UNAVAILABLE

Estimate queue
L x 1 - minRTT/sampleRtt

limit + 1

limit - 1

The loop never reads a capacity setting: it infers queueing from the gap between the current sample and the best latency it has seen. Source: Netflix/concurrency-limits README, checked 2026-09-29.
Diagram source

The registry that survived

Eureka is the only 2016 network component still described in the present tense by its own README, which says it "plays a critical role in Netflix mid-tier infra". It also says support is "Community-driven mostly". Both are true at once, and the incidents in section 4 are what that combination produces.

Evidence: Netflix/eureka README

The gateway that survived

Zuul is an L7 gateway you deploy and operate, not a dependency you link. It is the single component that spans the whole decade without a status notice, and the only one whose policy can be changed without touching a calling application.

Evidence: Netflix/zuul README

What is being built instead

The new in-house work visible in 2026 is a Java toolchain: jig resolves and assembles Java modules, and describes itself as "currently in preview". After a decade of moving decisions out of libraries, the fresh investment is in how libraries are built and shipped rather than in what they decide.

Evidence: Netflix/jig README

03

The decisions that matter

Five forks, each with the reason its owner published and the condition that would have sent them the other way.

Decision: should a concurrency threshold be configured or inferred?

Chosen
  • Inferred per node, by a delay-based control loop
  • Because a configured value "quickly goes out of date" once the fleet autoscales
  • Because the full set of binding resources cannot be enumerated
Rejected
  • A tuned threshold per service, the Hystrix model
  • Lost on maintenance: keeping thousands of thresholds correct requires operators to "fully understand the hardware services run on and coordinate how they scale"
Flips when
  • Your capacity is fixed and your dependency set is stable, so the number stays right once set
  • Or your traffic is bimodal enough that a latency-driven loop cannot tell a shape change from an overload, which is the unanswered question in section 4

Decision: does network policy ship as a library or as an interceptor?

Chosen
  • gRPC with load-balancing and discovery interceptors
  • Stated reasons: "multi-language support and better extensibility/composability through request interceptors"
Rejected
  • Continuing Ribbon, a Java client library
  • Lost because a single-language library cannot serve a fleet that has stopped being single-language
Flips when
  • One runtime, one language, one deployment train. Then a library is cheaper, and an interceptor layer is ceremony

Decision: how do you retire a library that thousands of builds already depend on?

Chosen
  • Make the dependency harmless rather than absent. Servo shipped a no-op registry as the default "to minimize the overhead for legacy apps that still happen to have some usage"
Rejected
  • Waiting for every consumer to remove the dependency
  • Lost because a transitive dependency has no owner to ask
Flips when
  • The library holds state or a contract that a no-op would silently corrupt. A metrics library can become a no-op; a serialiser cannot

The other two forks are about exits rather than mechanisms, and the repositories record three different ones with three different costs. Curator was donated: "Curator has moved to Apache. The Netflix Curator project will remain to hold Netflix extensions to Curator." Vector was retired into somebody else's product, and the note is unusually candid about why, saying the team "decided to lean into the Grafana stack" because "Grafana is widely used, well supported, and has an extensible framework". Conductor took the third route: on 13 December 2023 Netflix discontinued maintenance of the open-source project to realign on its internal fork. Each of those is the right answer under a different condition, and the condition is the thing worth carrying away.

DecisionChosenRejectedBecauseFlips whenEvidence
Circuit breakerAdaptive limiterTuned thresholdsConfigured value goes stale as the fleet scalesFixed capacity, stable dependenciesHystrix status notice
Load balancinggRPC interceptorsJava client libraryMulti-language fleet, composabilitySingle-language fleetRibbon status notice
Retiring a dependencyDefault to a no-opWait for removalTransitive dependencies have no ownerThe library holds a contract, not a side effectServo README
Exit route, standards existDonate upstreamKeep maintainingAn equivalent became the industry defaultNo credible standard exists yetCurator README
Exit route, fork divergedDiscontinue the public projectReconcile the forksInternal and public copies stopped agreeingThe divergence is small enough to mergeConductor archive notice, 2023-12-13
Exit route, suite of toolsDisperse into the delivery platformKeep the suiteEach function belonged to a different ownerThe functions genuinely share an operatorSimian Army retirement notice

Figure 5 · Where to put a policy, derived from what survived

yes

no

no

yes

yes

no

Does a human
keep a value correct?

Replace the value
with a control loop

Must it change
faster than a
fleet redeploy?

A library is fine:
formats, contracts,
instrumentation

Is every caller
the same runtime?

Interceptor shipped
with the RPC stack

A deployed service:
gateway or proxy

yes

no

no

yes

yes

no

Does a human
keep a value correct?

Replace the value
with a control loop

Must it change
faster than a
fleet redeploy?

A library is fine:
formats, contracts,
instrumentation

Is every caller
the same runtime?

Interceptor shipped
with the RPC stack

A deployed service:
gateway or proxy

The branch that decides everything is the first one: if a human has to keep a number correct, the number is the problem, not its location. Derived by the author from the reasons given in the Hystrix, Ribbon and Servo status notices.
Diagram source

One decision in this table is easy to mis-read, and the repositories are clear enough to settle it. Retiring Hystrix was not a judgement that circuit breaking was wrong. The notice says Netflix would "continue using Hystrix for existing applications", and points new internal projects at resilience4j. What was abandoned was the tuned threshold as an interface, not the bulkhead as an idea. An architect taking a lesson from this should not remove their circuit breakers. They should ask which of their thresholds a human is expected to keep correct, and how that human would find out they had stopped being correct.

04

What broke in production

Four incidents, in two classes, plus the class nobody has published.

How to grade these

Netflix publishes no post-incident reviews in this corpus. The four entries below are operator-filed reports in public issue trackers, which is a weaker artefact than a postmortem: no timeline discipline, no measured impact, and in three of the four cases no maintainer reply at all. They are included because they are the only production evidence that exists here, and because what they describe is consistent across independent reporters. Treat each as one practitioner's account, not as a finding.

Incident report

The refresh loop died, and nothing said so for four days

AssumptionA discovery client that cannot refresh will fail loudly, or at least fail in a way an alert catches.
What happenedThe client's scheduler exhausted a two-thread pool and logged "Task java.util.concurrent.FutureTask rejected from java.util.concurrent.ThreadPoolExecutor[Running, pool size = 2, active threads = 2, queued tasks = 0]" as a warning. The instance kept serving from its last good cache. Four days later a dependency was redeployed, the cached addresses became wrong, and calls failed with "No route to host".
Blast radiusOne pod, eureka-client 1.10.17 with Spring Cloud 3.1.2, resolved by a manual restart. Detection gap of four days, bounded by an unrelated deployment rather than by monitoring.
FixNone upstream. The reporter notes the behaviour reproduces across versions and references an earlier issue on the same mechanism.
Design ruleAny cache that is refreshed by a background loop needs the loop's liveness as a first-class signal, exported as an age, not as a log line. The failure is invisible until a second, unrelated change makes the stale data wrong, so the time between cause and symptom is set by somebody else's deploy cadence.
Incident report

The registry went away and every client kept answering confidently

AssumptionWhen the registry is unreachable, callers will find out.
What happenedWith the server down, the client could not update its local cache, and lookups kept returning instances: "When Eureka Server Down,EurekaClient can't change locat cache. So when I call DiscoveryClient getInstace,It still get application instance."
Blast radiusNot quantified by the reporter. Closed with a question label and no maintainer reply visible.
FixNone. This is the intended availability-favouring behaviour of the design, which is exactly what makes it dangerous to a caller who does not know that.
Design ruleServing stale data during a partition is a legitimate choice, but the staleness must be in the answer. A lookup that returns an address should be able to return how old that address is, or every caller inherits a decision they never made.
Incident report

The thing protecting the service allocates without a bound

AssumptionA load shedder fails safe, because refusing work is its whole purpose.
What happenedThe adaptive limiter's blocking executor defaults to a cached thread pool with no upper bound on thread creation. If the limiter is misconfigured or fails during a burst, the pool spawns a thread per request and the process reaches "java.lang.OutOfMemoryError: unable to create new native thread".
Blast radiusReported as a design hazard rather than an observed outage, opened 9 January 2026, no maintainer reply at the time of checking.
FixProposed by the reporter, not adopted: a bounded pool with an explicit rejection policy, or requiring the caller to supply an executor.
Design ruleEvery overload-control mechanism has a resource it consumes while deciding. Bound that resource independently of the mechanism, because the failure mode you care about is the one where the mechanism is what broke.
Incident report

The limiter's accounting leaks on the rejection path

AssumptionIn-flight counts are decremented on every path, including the ones that fail early.
What happenedWhen the gRPC executor rejects a task, the interceptor's in-flight accounting is not released. The operator's description of the symptom is the useful part: "the qps just dies after sometime".
Blast radiusThroughput decays to zero over time in the reporter's application. Opened 2 November 2023, still open, no reply visible.
FixNone published.
Design ruleA counter that gates admission is a resource with a release path, and the release path that matters is the abnormal one. Test the limiter by making its downstream reject, not by making it slow.

Figure 4 · Why the four-day detection gap was four days

DependencyApplicationLocal registry cacheRefresh loopDependencyApplicationLocal registry cacheRefresh looploop is dead, cache is frozenpool rejects task, logsWARNlookupaddresses, no age attachedcall succeeds for four daysredeployed to newaddressescall to stale addressNo route to host
DependencyApplicationLocal registry cacheRefresh loopDependencyApplicationLocal registry cacheRefresh looploop is dead, cache is frozenpool rejects task, logsWARNlookupaddresses, no age attachedcall succeeds for four daysredeployed to newaddressescall to stale addressNo route to host
Nothing fails at the moment of failure: the cache keeps answering, and the clock to the symptom is started by a deploy in a different team. Reconstructed from the timeline in Netflix/eureka issue 1510.
Diagram source

Grouping those four gives two classes. The first two are silent staleness: a background loop stops, the data it maintains keeps being served, and the failure only surfaces when something else changes. The second two are the control mechanism as a failure domain: the limiter allocates without a bound, and its accounting leaks on the path that runs precisely when the system is under stress. Both classes share a property worth naming, because it explains why they are so hard to catch: the symptom appears at a time chosen by an unrelated event, so there is no correlation for an operator to find.

There is a third class, and the honest statement about it is that nobody has published it. Issue 171 on the concurrency-limits tracker asks the question that decides whether the adaptive approach is sound: "How does this differentiate between a dependent service getting slower and standard too much concurrency impacting latency?" The reporter points out that during a dependency failure, latency spikes for a reason that reducing concurrency will not fix, and suggests adding a CPU or blocked-thread signal to disambiguate. It was opened on 27 July 2021 and has no reply. That is not evidence the mechanism is wrong. It is evidence that the most important operational question about the thing Netflix replaced Hystrix with has been open in public for five years, and an architect adopting it should plan to answer it themselves with a second signal that is not latency.

05

Numbers you can plan against

There is no latency, throughput or cost figure in this corpus. What there is instead is dates, and dates turn out to be the more useful number here.

MetricValueAtContextAs ofSource
Last commit touching Hystrix core source2021-11-30NetflixA typo fix in a class name, not a behaviour change2026-09-29commit history
Newest commit on the Hystrix repository2025-12-17NetflixAn org-wide CI change, four years after the last source change2026-09-29pull request 2115
Last commit touching Ribbon load-balancer source2021-03-03Netflix"Upgrade to modern gradle and nebula", a build change2026-09-29commit history
Repositories receiving the same CI commit on one day4NetflixHystrix, Ribbon, Servo and Governator, all deprecated, all touched 17 Dec 20252026-09-29servo, governator
Last commit touching concurrency-limits core source2026-01-12Netflix"Add time unit to Limit#onSample", a real change: the successor is maintained2026-09-29commit history
Age of the unanswered design question on that project5 yearsNetflixIssue 171, opened 2021-07-27, no reply2026-09-29issue 171
Longest time to close an outside pull request unmerged6.6 yearsNetflix"add deadline limiter", opened 2019-11-15, closed 2026-06-17. Derived by subtraction2026-09-29closed-unmerged list
Capabilities advertised by Spring Cloud Netflix8 then 2Spring1.4.x branch against 3.0.x and the current branch2026-09-291.4.x, main
Detection gap in the one detailed incident report4 daysOperatorBetween the refresh loop dying and the stale addresses mattering2026-09-29issue 1510
Date the public Conductor project was discontinued2023-12-13NetflixStated reason: realignment on the internal fork2026-09-29archive notice
Read these carefully

Every figure above is measured from a repository, with one exception: the 6.6-year pull-request age is derived by subtracting the two dates the listing shows. Nothing here is a vendor claim, because there is no vendor material in this corpus. The figures an architect would actually want, meaning the latency cost of an interceptor, the throughput recovered by an adaptive limiter against a tuned one, or the operational cost of running a gateway rather than a library, are unknown here: none appears in any repository reached, and the engineering blog that would carry them was unreachable. If you are making a case on those grounds, you will have to measure it yourself, and section 7 is where to start.

The two commit-date rows are the ones to carry into your own work, because together they break a method that looks reliable. A repository's front page shows its newest commit, and on 17 December 2025 one Netflix engineer pushed the same change, "Update Github Actions to use latest NetflixOSS recommendations", into Hystrix, Ribbon, Servo and Governator. All four are deprecated. All four now show a 2025 date to anyone judging liveness by the obvious signal. The real dates are 2021 and earlier, and you only see them by asking the commit history for a source path rather than for the repository.

06

The evidence wall

Every source behind this page, graded. There are no blogs, talks, papers or vendor case studies in it, and section 1 says why.

Decision record Netflix2018-11

Hystrix status notice

Declares maintenance mode and names the direction: "more adaptive implementations that react to an application's real time performance rather than pre-configured settings". New internal projects are pointed at resilience4j.

Carry forwardThe thing abandoned was the tuned threshold as an interface, not the bulkhead as a pattern.
github.com/Netflix/Hystrix
Source Netflix2018-11-09

Hystrix issue 1891, "Re-release Hystrix 1.5.11"

Netflix ran 1.5.11 internally while 1.5.13 was public, citing "issues and instabilities" in the newer version. The final 1.5.18 release realigns Maven Central with what Netflix actually ran.

Carry forwardThe published version of a company's library is not evidence of what that company runs.
github.com/Netflix/Hystrix/issues/1891
Decision record Netflixc. 2016

Ribbon, "Project Status: On Maintenance"

Names gRPC as the destination and gives two reasons: "multi-language support and better extensibility/composability through request interceptors". Also lists which of its own modules are "not used" internally.

Carry forwardA single-language client library has a shelf life set by how long the fleet stays single-language.
github.com/Netflix/ribbon
Source Netflix2026-01

concurrency-limits README and core history

The argument against configured limits, the Little's Law framing, and the Vegas and Gradient2 algorithms. Core source last changed 12 January 2026, so this is a live project.

Carry forwardThe mechanism concedes that you cannot enumerate your own bottlenecks, and designs around that rather than against it.
github.com/Netflix/concurrency-limits
Source Netflix2021-07-27

concurrency-limits issue 171

Asks how a latency-driven limiter tells a slow dependency from genuine self-overload, and proposes a CPU or blocked-thread signal. No reply in five years.

Carry forwardAdopt the adaptive limiter, then add the second signal yourself; the project has not.
github.com/Netflix/concurrency-limits/issues/171
Incident report Operator2026-01-09

concurrency-limits issue 231, unbounded thread creation

The blocking executor defaults to an unbounded cached thread pool, so a limiter failure during a burst yields "unable to create new native thread" rather than shedding.

Carry forwardBound the resources your overload control consumes, separately from the control itself.
github.com/Netflix/concurrency-limits/issues/231
Incident report Operator2023-11-02

concurrency-limits issue 190, in-flight leak

In-flight accounting is not released when the gRPC executor rejects a task. Reported symptom: "the qps just dies after sometime". Still open.

Carry forwardExercise the admission counter's abnormal release path, which is the one that runs under stress.
github.com/Netflix/concurrency-limits/issues/190
Incident report Operator2023-08-07

Eureka issue 1510, refresh thread pool exhausted

The most detailed production account here: a two-thread scheduler rejects its own task, the client serves a stale registry for four days, and the failure surfaces only when a dependency redeploys.

Carry forwardExport cache age as a metric; a warning log is not a detection mechanism.
github.com/Netflix/eureka/issues/1510
Incident report Operator2020-11-06

Eureka issue 1362, stale cache when the server is down

With the registry unreachable the client keeps returning instances from its cache. Closed as a question with no maintainer reply.

Carry forwardIf a lookup can return stale data, it must be able to return how stale.
github.com/Netflix/eureka/issues/1362
Decision record Netflixn/d

Servo deprecation and the no-op registry

Deprecated in favour of Spectator. Since 0.13.0 the default monitor registry is a no-op "to minimize the overhead for legacy apps that still happen to have some usage of Servo".

Carry forwardYou cannot uninstall a transitive dependency, but you can make it cost nothing.
github.com/Netflix/servo
Decision record Netflixn/d

Simian Army, "PROJECT STATUS: RETIRED"

Dissolved rather than replaced: Chaos Monkey became a standalone service, Swabbie took over Janitor Monkey's work, Conformity Monkey moved into Spinnaker.

Carry forwardWhen a tool suite retires, look at which platform absorbed each function; that names the real owner.
github.com/Netflix/SimianArmy
Decision record Netflixn/d

Vector retirement notice

Netflix retired its own performance-monitoring front end and contributed a Grafana data source to PCP: "We have decided to lean into the Grafana stack. Grafana is widely used, well supported, and has an extensible framework."

Carry forwardThe build-versus-adopt decision reverses when an ecosystem tool grows the extension point you built your own tool for.
github.com/Netflix/vector
Decision record Netflixn/d

Curator, moved to Apache

"Curator has moved to Apache. The Netflix Curator project will remain to hold Netflix extensions to Curator." The donation exit.

Carry forwardDonating upstream and keeping a thin extension repository is the cheapest exit, and the only one that leaves users unharmed.
github.com/Netflix/curator
Decision record Netflix2023-12-13

Conductor discontinuation notice

"Effective December 13, 2023, Netflix will discontinue maintenance of Conductor OSS on GitHub." The stated motivation is realigning resources on the internal fork.

Carry forwardWhen the internal and public copies of a system stop agreeing, the public copy is the one that gets dropped.
github.com/Netflix/conductor
Source Netflix2025-12-17

Hystrix pull request 2115, the CI commit

A publishing configuration change merged into four deprecated repositories on one day. It is why those repositories show recent activity.

Carry forwardA repository's newest commit is not a liveness signal; ask the history for a source path instead.
github.com/Netflix/Hystrix/pull/2115
Source Netflix2021-11-30

Hystrix core source history

The last commit touching hystrix-core/src is a typo fix in a class name, dated 30 November 2021.

Carry forwardThe functional death date of a dependency is the last change to its source, and it is usually years before the archive banner.
github.com/Netflix/Hystrix/commits/master/hystrix-core/src
Source Netflix2018-2026

concurrency-limits closed-unmerged pull requests

"add deadline limiter" was open from 15 November 2019 to 17 June 2026; "Add counters for partitions" from December 2020 to September 2025.

Carry forwardTime-to-decision on outside contributions is the number that tells you whether you can influence a dependency, and years means no.
github.com/Netflix/concurrency-limits/pulls
Source Springbranch 1.4.x

spring-cloud-netflix README, 1.4.x branch

The adopted stack as the ecosystem saw it: Eureka, Hystrix and its dashboard, Feign, Ribbon, the Archaius bridge, Zuul filters.

Carry forwardReading an integration project's old branches is the cheapest way to date what an ecosystem actually depended on.
raw.githubusercontent.com/spring-cloud/spring-cloud-netflix/1.4.x/README.adoc
Decision record Springbranch 2.2.x

spring-cloud-netflix reference docs, Ribbon migration note

"Since Ribbon load-balancer is now in maintenance mode, we suggest switching to using the Spring Cloud LoadBalancer." The downstream project built its own replacement.

Carry forwardWhen you adopt another organisation's library, you inherit their deprecation calendar, and your migration date is chosen by them.
raw.githubusercontent.com/spring-cloud/spring-cloud-netflix/2.2.x/docs/.../spring-cloud-netflix.adoc
Source resilience4jn/d

resilience4j README

Named by Netflix as the successor for new projects. The copyright line lists individual maintainers and there is no sponsorship statement.

Carry forwardA successor named in a deprecation notice is a recommendation, not a support commitment; check who is actually behind it.
github.com/resilience4j/resilience4j
Source Netflix2026-09-29

Netflix organisation repository listing

Sorted by last update, the September 2026 top of the list is Atlas, Spectator, EVCache, DGS, Hollow, Genie, Metaflow, Zuul, Mantis, Metacat, Maestro and a new Java toolchain. None of the retired network libraries appears.

Carry forwardSorting an organisation's repositories by update time, then asking what each one decides, separates the estate faster than reading any of them.
github.com/orgs/Netflix/repositories
Source Netflixn/d

Zuul README

Deployed as a service, still committed to in September 2026, and the only 2016-era network component with no status notice.

Carry forwardThe component that survived the decade is the one whose policy changes without a caller redeploying.
github.com/Netflix/zuul

The mix is deliberately narrow and the narrowness is the weakness of this page. There are four incident reports and none of them is a postmortem. There is no talk, no paper and no independent measurement. What the corpus is good for is chronology and stated reasoning, because deprecation notices and archive banners are written once, dated, and rarely revised. What it cannot support is any claim about how well the replacements perform, and this guide makes none.

07

Build a miniature, then productionise it

Six rungs. The first three land the argument, the last three produce the evidence this corpus is missing.

Set a threshold, then make it wrong

Put a fixed concurrency limit in front of a toy service, load test it, and set the limit at seventy-five percent of the tipping point. Then double the instance size and run the same test.

Done when: the limit you chose is rejecting traffic the service could have served, and you can state by how much.  Teaches: why the concurrency-limits README calls a configured value one that "quickly goes out of date".

Implement Vegas in fifty lines

Track minRTT, sample latency per request, estimate the queue as L * (1 - minRTT/sampleRtt), and move the limit by one per window.

Done when: the limit tracks a step change in instance size without you touching a setting.  Teaches: that the control loop is small, and that all its difficulty is in what it measures.

Break the loop's assumption on purpose

Make a downstream dependency slow while offered load stays flat. Watch the limiter reduce concurrency for a condition that reducing concurrency does not fix.

Done when: you have reproduced the scenario in issue 171 and measured the throughput you lost.  Teaches: why a second signal that is not latency is a requirement rather than a refinement.

Kill a refresh loop silently

Run a service with a cached service-discovery result. Starve the refresh thread pool, leave the cache serving, then redeploy the dependency to new addresses.

Done when: you can state your detection time, and it is longer than you expected.  Teaches: the failure in Eureka issue 1510, which is a property of the pattern rather than of that implementation.

Move one policy out of the process

Take the retry or timeout rule from your service's code and move it into an interceptor shipped with your RPC stack, or into a gateway. Change it without rebuilding the service.

Done when: you have changed the policy in production without a service redeploy, and you can name every consumer the change reached.  Teaches: what Netflix bought with the Ribbon-to-gRPC move, and what it cost in new failure modes.

Excavate your own dependency list

For each third-party library in your critical path, find the last commit to its source directory, the age of its oldest unanswered issue, and the time-to-close on its last ten outside pull requests.

Done when: you can rank your dependencies by how long a fix would take to reach you.  Teaches: that a dependency's risk is a schedule you do not control, which is the lesson of this whole page.

08

Keep hunting

The URL shapes that produced this page. They work on any organisation's repositories and need no search engine, which is what made them usable here.

Dating a dependency honestly

  • github.com/ORG/REPO/commits/BRANCH/PATH-TO-src
  • github.com/orgs/ORG/repositories?type=source&sort=updated
  • raw.githubusercontent.com/ORG/REPO/BRANCH/README.md

Finding the argument and the refusal

  • github.com/ORG/REPO/pulls?q=is:pr+is:closed+is:unmerged
  • github.com/ORG/REPO/issues?q=is:issue+outage
  • github.com/ORG/REPO/issues?q=is:issue+sort:reactions-%2B1-desc

The move worth keeping is the branch diff on an integration project. Fetching the same README.adoc from the 1.4.x, 2.2.x, 3.0.x and current branches of spring-cloud-netflix dates the ecosystem's dependency on each Netflix component without reading a single release note. Any project that integrates one organisation's stack into another's has this property, and its old branches stay fetchable long after the announcements are gone.

09

References

  1. Netflix, Hystrix status notice GitHub README, notice dated November 2018, checked 2026-09-29.
  2. Netflix, "Re-release Hystrix 1.5.11", issue 1891 GitHub issue, 9 November 2018, checked 2026-09-29.
  3. Netflix, "Update Github Actions to use latest NetflixOSS recommendations", pull request 2115 GitHub pull request, 17 December 2025, checked 2026-09-29.
  4. Netflix, Hystrix core source commit history GitHub commit history, checked 2026-09-29.
  5. Netflix, Ribbon, "Project Status: On Maintenance" GitHub README, checked 2026-09-29.
  6. Netflix, Ribbon load-balancer source commit history GitHub commit history, checked 2026-09-29.
  7. Netflix, concurrency-limits GitHub README, checked 2026-09-29.
  8. Netflix, concurrency-limits core source commit history GitHub commit history, latest 12 January 2026, checked 2026-09-29.
  9. Netflix, concurrency-limits issue 171 GitHub issue, 27 July 2021, checked 2026-09-29.
  10. Netflix, concurrency-limits issue 190 GitHub issue, 2 November 2023, checked 2026-09-29.
  11. Netflix, concurrency-limits issue 231 GitHub issue, 9 January 2026, checked 2026-09-29.
  12. Netflix, concurrency-limits closed-unmerged pull requests GitHub pull request listing, checked 2026-09-29.
  13. Netflix, Eureka GitHub README, checked 2026-09-29.
  14. Netflix, Eureka issue 1362 GitHub issue, 6 November 2020, checked 2026-09-29.
  15. Netflix, Eureka issue 1510 GitHub issue, 7 August 2023, checked 2026-09-29.
  16. Netflix, Servo deprecation notice GitHub README, checked 2026-09-29.
  17. Netflix, Servo commit history GitHub commit history, checked 2026-09-29.
  18. Netflix, Governator commit history GitHub commit history, checked 2026-09-29.
  19. Netflix, Simian Army retirement notice GitHub README, checked 2026-09-29.
  20. Netflix, Vector retirement notice GitHub README, checked 2026-09-29.
  21. Netflix, Curator GitHub README, checked 2026-09-29.
  22. Netflix, Conductor discontinuation notice GitHub repository, notice dated 13 December 2023, checked 2026-09-29.
  23. Netflix, Titus archival notice GitHub README, checked 2026-09-29.
  24. Netflix, Archaius GitHub README, checked 2026-09-29.
  25. Netflix, Zuul GitHub README, checked 2026-09-29.
  26. Netflix, Maestro GitHub README, checked 2026-09-29.
  27. Netflix, jig GitHub README, preview, checked 2026-09-29.
  28. Netflix, DGS framework releases GitHub releases, checked 2026-09-29.
  29. Netflix, Falcor commit history GitHub commit history, checked 2026-09-29.
  30. Netflix, organisation source repositories by last update GitHub organisation listing, checked 2026-09-29.
  31. Spring, spring-cloud-netflix README, 1.4.x branch GitHub raw file, checked 2026-09-29.
  32. Spring, spring-cloud-netflix README, 3.0.x branch GitHub raw file, checked 2026-09-29.
  33. Spring, spring-cloud-netflix README, current branch GitHub raw file, checked 2026-09-29.
  34. Spring, spring-cloud-netflix reference documentation, 2.2.x branch GitHub raw file, checked 2026-09-29.
  35. resilience4j GitHub README, checked 2026-09-29.