Hystrix status notice
Declares maintenance mode and names the direction: "more adaptive implementations that react to an application's real time performance rather than pre-configured settings". New internal projects are pointed at resilience4j.
Between 2016 and 2026 almost every Netflix library that decided something about the network was retired, while the libraries that decide something about data survived. This guide reconstructs that decade from the repositories themselves, and turns it into a rule for deciding where a policy belongs in a system you are designing now.
A company changed its mind about where a rule should live. The record of it is in thirty repositories, and it is more specific than any retrospective.
Here is the problem without a product name attached to it. A large fleet of processes has to obey rules about the network: which instance to call, how many calls may be in flight, when to stop calling, what to do when the answer is slow. Those rules have to change faster than the processes do, because traffic shifts, capacity changes and dependencies fail on their own schedule. If a rule lives inside a library that every process links, then changing the rule means redeploying the fleet, and the number of teams who must agree to that is the real cost of the design. The question this guide answers is what happens to an architecture over ten years when that cost is paid repeatedly, and what the organisation does about it once it notices.
Netflix is the case worth excavating because its 2016 answer was the one the industry copied. Spring Cloud Netflix, which is how most organisations actually adopted it, advertised eight capabilities on its 1.4.x branch: discovery, an embedded discovery server, a circuit breaker and its dashboard, a declarative REST client, a client-side load balancer, a configuration bridge and a routing filter. Seven of the eight are rules about the network, and every one shipped as a Java library linked into the application process. On the current branch the same file lists two, and both are Eureka.
The session that researched this page reached three hosts: github.com,
raw.githubusercontent.com and gist.github.com. Every other host
was refused by a network egress policy, including Netflix's own engineering blog, the
conference and paper archives, and the trade press. So there is no blog post, talk, paper
or cost figure in this corpus, and Netflix's own narrative of these changes is absent.
What remains is the layer that is hardest to spin and easiest to date: deprecation
notices, archive banners, commit histories, release trains and issue threads. The guide is
written entirely from those, and every place where they run out is marked.
The finding that changed how this guide is organised is not that the libraries were retired. It is which ones. Netflix in September 2026 still ships Spectator, Hollow, the DGS GraphQL framework, Metaflow, Atlas, EVCache and Zuul, all updated within the last week of the organisation's repository listing. Hystrix, Ribbon, Archaius, Servo, Governator, the Simian Army, Conductor, Falcor and Titus are absent from the first thirty entries. Sort by what each thing decides and the split is clean: the survivors encode a data or instrumentation contract, the casualties encoded a rule about the network. Sections 2 and 3 test that against the stated reasons.
One related excavation already exists in this collection and asks a different question of the same kind of evidence: The half-life of in-house infrastructure reads a decade of Uber's repositories to classify how in-house systems exit, as donated, deprecated or drifted. This guide takes the exits as given and asks what replaced the thing, and what the replacements have in common. Where the two overlap, on how to read an archive banner and what a commit date is worth, section 8 corrects the earlier method rather than repeating it.
The 2016 shape, the 2026 shape, and the three places the policy went.
The 2016 architecture has an unusual property: there is no control plane. Discovery, load balancing, circuit breaking and configuration all live in the calling process, as objects configured by properties that ship with the application. The registry server and the dashboard exist, but they are data sources for in-process logic rather than places where a decision is made. That is why the design scaled so well at first and became expensive later. Every call is decided locally, which needs no coordination; every change to how calls are decided needs every caller to redeploy.
The 2026 shape, reconstructed from what Netflix retired and what it points at instead, puts the same four decisions in three different places. Load balancing and its associated cross-cutting concerns moved into RPC interceptors: the Ribbon README states the destination directly, that the team "started building an RPC solution on top of gRPC" for "multi-language support and better extensibility/composability through request interceptors". Circuit breaking moved into a control loop that sets its own threshold, described in the next paragraph. Routing and edge policy stayed in Zuul, which is not a library at all but a gateway you deploy, and which is one of the few 2016 components still being committed to in September 2026. Configuration, in the Archaius sense of a library that reads properties into typed objects, largely disappeared as a separate concern, because a control loop that infers its own limits does not need a property to read.
Worth being precise about what did and did not move, because the popular version of this story is wrong. Netflix did not move resilience out of the process into a sidecar. An interceptor runs in the same process as the caller, in the same language, on the same thread. What changed is not the location of the code but the location of the decision: an interceptor is shipped and versioned with the RPC stack rather than written per-application, and a control loop takes its setting from measurement rather than from a property file that somebody has to be right about. The code stayed in the application. The judgement left it.
The replacement for the circuit breaker deserves reading in the original, because its
README contains the clearest statement of the problem anyone involved wrote down. Netflix's
concurrency-limits project argues that operators think in requests per second, set a limit
below a measured tipping point, and then discover that "in large distributed systems that
auto-scale this value quickly goes out of date and the service falls over by becoming
non-responsive as it is unable to gracefully shed excess load". Its proposed alternative is
to reason in concurrent requests instead, via Little's Law, and then to refuse to configure
even that: the limit is estimated per node by treating it as a TCP congestion window and
running a delay-based algorithm over observed latency. The Vegas variant estimates the
queue as L * (1 - minRTT/sampleRtt) and moves the limit by one each sampling
window. The design concedes its own motivation in a single line: "For large and complex
distributed systems it's impossible to know all the hard resources."
Eureka is the only 2016 network component still described in the present tense by its own README, which says it "plays a critical role in Netflix mid-tier infra". It also says support is "Community-driven mostly". Both are true at once, and the incidents in section 4 are what that combination produces.
Evidence: Netflix/eureka README
Zuul is an L7 gateway you deploy and operate, not a dependency you link. It is the single component that spans the whole decade without a status notice, and the only one whose policy can be changed without touching a calling application.
Evidence: Netflix/zuul README
The new in-house work visible in 2026 is a Java toolchain: jig resolves
and assembles Java modules, and describes itself as "currently in preview". After a
decade of moving decisions out of libraries, the fresh investment is in how libraries
are built and shipped rather than in what they decide.
Evidence: Netflix/jig README
Five forks, each with the reason its owner published and the condition that would have sent them the other way.
The other two forks are about exits rather than mechanisms, and the repositories record three different ones with three different costs. Curator was donated: "Curator has moved to Apache. The Netflix Curator project will remain to hold Netflix extensions to Curator." Vector was retired into somebody else's product, and the note is unusually candid about why, saying the team "decided to lean into the Grafana stack" because "Grafana is widely used, well supported, and has an extensible framework". Conductor took the third route: on 13 December 2023 Netflix discontinued maintenance of the open-source project to realign on its internal fork. Each of those is the right answer under a different condition, and the condition is the thing worth carrying away.
| Decision | Chosen | Rejected | Because | Flips when | Evidence |
|---|---|---|---|---|---|
| Circuit breaker | Adaptive limiter | Tuned thresholds | Configured value goes stale as the fleet scales | Fixed capacity, stable dependencies | Hystrix status notice |
| Load balancing | gRPC interceptors | Java client library | Multi-language fleet, composability | Single-language fleet | Ribbon status notice |
| Retiring a dependency | Default to a no-op | Wait for removal | Transitive dependencies have no owner | The library holds a contract, not a side effect | Servo README |
| Exit route, standards exist | Donate upstream | Keep maintaining | An equivalent became the industry default | No credible standard exists yet | Curator README |
| Exit route, fork diverged | Discontinue the public project | Reconcile the forks | Internal and public copies stopped agreeing | The divergence is small enough to merge | Conductor archive notice, 2023-12-13 |
| Exit route, suite of tools | Disperse into the delivery platform | Keep the suite | Each function belonged to a different owner | The functions genuinely share an operator | Simian Army retirement notice |
One decision in this table is easy to mis-read, and the repositories are clear enough to settle it. Retiring Hystrix was not a judgement that circuit breaking was wrong. The notice says Netflix would "continue using Hystrix for existing applications", and points new internal projects at resilience4j. What was abandoned was the tuned threshold as an interface, not the bulkhead as an idea. An architect taking a lesson from this should not remove their circuit breakers. They should ask which of their thresholds a human is expected to keep correct, and how that human would find out they had stopped being correct.
Four incidents, in two classes, plus the class nobody has published.
Netflix publishes no post-incident reviews in this corpus. The four entries below are operator-filed reports in public issue trackers, which is a weaker artefact than a postmortem: no timeline discipline, no measured impact, and in three of the four cases no maintainer reply at all. They are included because they are the only production evidence that exists here, and because what they describe is consistent across independent reporters. Treat each as one practitioner's account, not as a finding.
Grouping those four gives two classes. The first two are silent staleness: a background loop stops, the data it maintains keeps being served, and the failure only surfaces when something else changes. The second two are the control mechanism as a failure domain: the limiter allocates without a bound, and its accounting leaks on the path that runs precisely when the system is under stress. Both classes share a property worth naming, because it explains why they are so hard to catch: the symptom appears at a time chosen by an unrelated event, so there is no correlation for an operator to find.
There is a third class, and the honest statement about it is that nobody has published it. Issue 171 on the concurrency-limits tracker asks the question that decides whether the adaptive approach is sound: "How does this differentiate between a dependent service getting slower and standard too much concurrency impacting latency?" The reporter points out that during a dependency failure, latency spikes for a reason that reducing concurrency will not fix, and suggests adding a CPU or blocked-thread signal to disambiguate. It was opened on 27 July 2021 and has no reply. That is not evidence the mechanism is wrong. It is evidence that the most important operational question about the thing Netflix replaced Hystrix with has been open in public for five years, and an architect adopting it should plan to answer it themselves with a second signal that is not latency.
There is no latency, throughput or cost figure in this corpus. What there is instead is dates, and dates turn out to be the more useful number here.
| Metric | Value | At | Context | As of | Source |
|---|---|---|---|---|---|
| Last commit touching Hystrix core source | 2021-11-30 | Netflix | A typo fix in a class name, not a behaviour change | 2026-09-29 | commit history |
| Newest commit on the Hystrix repository | 2025-12-17 | Netflix | An org-wide CI change, four years after the last source change | 2026-09-29 | pull request 2115 |
| Last commit touching Ribbon load-balancer source | 2021-03-03 | Netflix | "Upgrade to modern gradle and nebula", a build change | 2026-09-29 | commit history |
| Repositories receiving the same CI commit on one day | 4 | Netflix | Hystrix, Ribbon, Servo and Governator, all deprecated, all touched 17 Dec 2025 | 2026-09-29 | servo, governator |
| Last commit touching concurrency-limits core source | 2026-01-12 | Netflix | "Add time unit to Limit#onSample", a real change: the successor is maintained | 2026-09-29 | commit history |
| Age of the unanswered design question on that project | 5 years | Netflix | Issue 171, opened 2021-07-27, no reply | 2026-09-29 | issue 171 |
| Longest time to close an outside pull request unmerged | 6.6 years | Netflix | "add deadline limiter", opened 2019-11-15, closed 2026-06-17. Derived by subtraction | 2026-09-29 | closed-unmerged list |
| Capabilities advertised by Spring Cloud Netflix | 8 then 2 | Spring | 1.4.x branch against 3.0.x and the current branch | 2026-09-29 | 1.4.x, main |
| Detection gap in the one detailed incident report | 4 days | Operator | Between the refresh loop dying and the stale addresses mattering | 2026-09-29 | issue 1510 |
| Date the public Conductor project was discontinued | 2023-12-13 | Netflix | Stated reason: realignment on the internal fork | 2026-09-29 | archive notice |
Every figure above is measured from a repository, with one exception: the 6.6-year pull-request age is derived by subtracting the two dates the listing shows. Nothing here is a vendor claim, because there is no vendor material in this corpus. The figures an architect would actually want, meaning the latency cost of an interceptor, the throughput recovered by an adaptive limiter against a tuned one, or the operational cost of running a gateway rather than a library, are unknown here: none appears in any repository reached, and the engineering blog that would carry them was unreachable. If you are making a case on those grounds, you will have to measure it yourself, and section 7 is where to start.
The two commit-date rows are the ones to carry into your own work, because together they break a method that looks reliable. A repository's front page shows its newest commit, and on 17 December 2025 one Netflix engineer pushed the same change, "Update Github Actions to use latest NetflixOSS recommendations", into Hystrix, Ribbon, Servo and Governator. All four are deprecated. All four now show a 2025 date to anyone judging liveness by the obvious signal. The real dates are 2021 and earlier, and you only see them by asking the commit history for a source path rather than for the repository.
Every source behind this page, graded. There are no blogs, talks, papers or vendor case studies in it, and section 1 says why.
Declares maintenance mode and names the direction: "more adaptive implementations that react to an application's real time performance rather than pre-configured settings". New internal projects are pointed at resilience4j.
Netflix ran 1.5.11 internally while 1.5.13 was public, citing "issues and instabilities" in the newer version. The final 1.5.18 release realigns Maven Central with what Netflix actually ran.
Names gRPC as the destination and gives two reasons: "multi-language support and better extensibility/composability through request interceptors". Also lists which of its own modules are "not used" internally.
The argument against configured limits, the Little's Law framing, and the Vegas and Gradient2 algorithms. Core source last changed 12 January 2026, so this is a live project.
Asks how a latency-driven limiter tells a slow dependency from genuine self-overload, and proposes a CPU or blocked-thread signal. No reply in five years.
The blocking executor defaults to an unbounded cached thread pool, so a limiter failure during a burst yields "unable to create new native thread" rather than shedding.
In-flight accounting is not released when the gRPC executor rejects a task. Reported symptom: "the qps just dies after sometime". Still open.
The most detailed production account here: a two-thread scheduler rejects its own task, the client serves a stale registry for four days, and the failure surfaces only when a dependency redeploys.
With the registry unreachable the client keeps returning instances from its cache. Closed as a question with no maintainer reply.
Deprecated in favour of Spectator. Since 0.13.0 the default monitor registry is a no-op "to minimize the overhead for legacy apps that still happen to have some usage of Servo".
Dissolved rather than replaced: Chaos Monkey became a standalone service, Swabbie took over Janitor Monkey's work, Conformity Monkey moved into Spinnaker.
Netflix retired its own performance-monitoring front end and contributed a Grafana data source to PCP: "We have decided to lean into the Grafana stack. Grafana is widely used, well supported, and has an extensible framework."
"Curator has moved to Apache. The Netflix Curator project will remain to hold Netflix extensions to Curator." The donation exit.
"Effective December 13, 2023, Netflix will discontinue maintenance of Conductor OSS on GitHub." The stated motivation is realigning resources on the internal fork.
A publishing configuration change merged into four deprecated repositories on one day. It is why those repositories show recent activity.
The last commit touching hystrix-core/src is a typo fix in a class name,
dated 30 November 2021.
"add deadline limiter" was open from 15 November 2019 to 17 June 2026; "Add counters for partitions" from December 2020 to September 2025.
The adopted stack as the ecosystem saw it: Eureka, Hystrix and its dashboard, Feign, Ribbon, the Archaius bridge, Zuul filters.
"Since Ribbon load-balancer is now in maintenance mode, we suggest switching to using the Spring Cloud LoadBalancer." The downstream project built its own replacement.
Named by Netflix as the successor for new projects. The copyright line lists individual maintainers and there is no sponsorship statement.
Sorted by last update, the September 2026 top of the list is Atlas, Spectator, EVCache, DGS, Hollow, Genie, Metaflow, Zuul, Mantis, Metacat, Maestro and a new Java toolchain. None of the retired network libraries appears.
Deployed as a service, still committed to in September 2026, and the only 2016-era network component with no status notice.
The mix is deliberately narrow and the narrowness is the weakness of this page. There are four incident reports and none of them is a postmortem. There is no talk, no paper and no independent measurement. What the corpus is good for is chronology and stated reasoning, because deprecation notices and archive banners are written once, dated, and rarely revised. What it cannot support is any claim about how well the replacements perform, and this guide makes none.
Six rungs. The first three land the argument, the last three produce the evidence this corpus is missing.
Put a fixed concurrency limit in front of a toy service, load test it, and set the limit at seventy-five percent of the tipping point. Then double the instance size and run the same test.
Done when: the limit you chose is rejecting traffic the service could have served, and you can state by how much. Teaches: why the concurrency-limits README calls a configured value one that "quickly goes out of date".
Track minRTT, sample latency per request, estimate the queue as
L * (1 - minRTT/sampleRtt), and move the limit by one per window.
Done when: the limit tracks a step change in instance size without you touching a setting. Teaches: that the control loop is small, and that all its difficulty is in what it measures.
Make a downstream dependency slow while offered load stays flat. Watch the limiter reduce concurrency for a condition that reducing concurrency does not fix.
Done when: you have reproduced the scenario in issue 171 and measured the throughput you lost. Teaches: why a second signal that is not latency is a requirement rather than a refinement.
Run a service with a cached service-discovery result. Starve the refresh thread pool, leave the cache serving, then redeploy the dependency to new addresses.
Done when: you can state your detection time, and it is longer than you expected. Teaches: the failure in Eureka issue 1510, which is a property of the pattern rather than of that implementation.
Take the retry or timeout rule from your service's code and move it into an interceptor shipped with your RPC stack, or into a gateway. Change it without rebuilding the service.
Done when: you have changed the policy in production without a service redeploy, and you can name every consumer the change reached. Teaches: what Netflix bought with the Ribbon-to-gRPC move, and what it cost in new failure modes.
For each third-party library in your critical path, find the last commit to its source directory, the age of its oldest unanswered issue, and the time-to-close on its last ten outside pull requests.
Done when: you can rank your dependencies by how long a fix would take to reach you. Teaches: that a dependency's risk is a schedule you do not control, which is the lesson of this whole page.
The URL shapes that produced this page. They work on any organisation's repositories and need no search engine, which is what made them usable here.
github.com/ORG/REPO/commits/BRANCH/PATH-TO-srcgithub.com/orgs/ORG/repositories?type=source&sort=updatedraw.githubusercontent.com/ORG/REPO/BRANCH/README.mdgithub.com/ORG/REPO/pulls?q=is:pr+is:closed+is:unmergedgithub.com/ORG/REPO/issues?q=is:issue+outagegithub.com/ORG/REPO/issues?q=is:issue+sort:reactions-%2B1-descThe move worth keeping is the branch diff on an integration project. Fetching the same
README.adoc from the 1.4.x, 2.2.x, 3.0.x and current branches of
spring-cloud-netflix dates the ecosystem's dependency on each Netflix component without
reading a single release note. Any project that integrates one organisation's stack into
another's has this property, and its old branches stay fetchable long after the
announcements are gone.