Cloudflare, 2016–2026  / field guide
Practitioner field guide · 2026-10-08 · Platform & Infrastructure

Hardening the wrong plane: ten years of Cloudflare's architecture

Between 2016 and 2026 Cloudflare rebuilt almost every component of its data plane, and made its configuration distribution faster at every step. Reconstructed from eight of its own incident reports, its engineering accounts and the issue threads in its open-source proxy, this guide traces which of those rebuilds paid, and why the outages kept arriving through the one path that never got release engineering: the push that carries data rather than code.

49 ledger rows 8 published incidents 15 source hosts Evidence through October 2026 Read: 40 min
01

The territory

A company runs one fleet of identical machines in hundreds of cities. It must be able to change what those machines do, globally, within seconds. Ten years of evidence on what that capability costs.

40M+
requests per second served by the proxy, by its own README
256ms
p99 to replicate a configuration write to 300+ locations
28%
of HTTP traffic broken by one config push, 5 December 2025
6 of 8
published incidents here began with a change to configuration, not code

State the problem without the company's name in it and it becomes general. You operate a large number of identical machines. The machines run code, which you release carefully, and they read configuration, which decides what the code does to each customer's traffic. Code moves slowly because releasing it is expensive. Configuration moves fast because the whole point of it is to move fast: a new attack signature is worthless if it takes a day to reach the edge. So you build a distribution system that puts a new value on every machine on earth in under a second, and you have thereby built the most powerful single-action blast radius in your architecture, and it is not covered by any of your release engineering.

Cloudflare is the best-documented instance of this in public: it publishes long incident reports with timelines and named mechanisms, and its architecture is unusually uniform, because every server in every city can serve any service on any of its addresses, so there are no natural partitions to limit a bad value's reach. The ten-year arc is easy to summarise and hard to accept. The data plane was rebuilt component by component into something markedly faster and safer. The configuration plane was rebuilt too, and each rebuild made it faster: in 2015 the team replaced the config store because what "worked for the first 25 cities began to show its age as the network passed 100", and in 2026 the successor to that replacement reaches over 300 locations with a p99 of 256 ms. Nothing in the public record describes a corresponding increase in the number of gates a configuration change has to pass.

The finding that surprised me

On 18 November 2025 the same oversized data file hit two generations of proxy, and the new one failed worse. Cloudflare's report says customers on the Rust proxy, FL2, "saw 5xx errors for most bot-protected traffic, while FL customers saw bot scores default to zero instead". The old unstructured Lua served wrong answers; the new memory-safe Rust called unwrap() on a failing result and panicked. Memory safety is not failure safety, and a rewrite that tightens correctness without moving validation to the ingestion boundary converts data errors into availability errors.

Who else has this problem: anyone running a feature-flag service, a WAF ruleset, a threat-intelligence feed, a routing table, a pricing table, a model artefact, or any machine-generated file that a fleet reads at request time. The specifics below are Cloudflare's. The shape is not.

Figure 1 · Two tracks, ten years

2015-2020Quicksilver replacesKyoto TycoonWorkers ship on V8isolates (2018)Unimog, eBPF/XDPload balancer (2020)WAF rule exhaustsCPU worldwide(2019)Backbone configdraws all traffic toAtlanta (2020)2021-2023Pingora replacesNGINX upstreamOxy becomes theinternal proxyframeworkBGP policy ordertakes 19 sites offline(2022)Utility fault takes thecontrol plane down(2023)2024-2026Topaz verifies DNSconfig formallyFL2 replaces NGINXand Lua with RustQuicksilver v2 andsub-second globalwritesWorkers KV storagedependency fails(2025)1.1.1.1 withdrawn by afive-week-old error(2025)Feature file panicsthe proxy (18 Nov2025)Config push breaks28% of traffic (5 Dec2025)Data plane rebuilt, configuration plane accelerated
2015-2020Quicksilver replacesKyoto TycoonWorkers ship on V8isolates (2018)Unimog, eBPF/XDPload balancer (2020)WAF rule exhaustsCPU worldwide(2019)Backbone configdraws all traffic toAtlanta (2020)2021-2023Pingora replacesNGINX upstreamOxy becomes theinternal proxyframeworkBGP policy ordertakes 19 sites offline(2022)Utility fault takes thecontrol plane down(2023)2024-2026Topaz verifies DNSconfig formallyFL2 replaces NGINXand Lua with RustQuicksilver v2 andsub-second globalwritesWorkers KV storagedependency fails(2025)1.1.1.1 withdrawn by afive-week-old error(2025)Feature file panicsthe proxy (18 Nov2025)Config push breaks28% of traffic (5 Dec2025)Data plane rebuilt, configuration plane accelerated
Within each era the rebuilds are listed first and the incidents after them. Read down the right-hand column: the configuration path's own rebuilds made it faster rather than more gated, and the incidents kept arriving through it. Compiled from the postmortems and engineering accounts in section 6.
Diagram source
Scope, and how this was researched

This guide covers the request path, the compute platform and the configuration distribution path, 2016 to 2026. It does not cover Cloudflare's security products as products, its commercial performance, the Workers developer experience, or any comparison with other content delivery networks. Claims about internal systems stop where the public record does: Oxy's internals, Quicksilver v2's storage engine and the current FL1-to-FL2 traffic split are all unpublished, and are marked as open below.

On method: this session's network policy permitted direct fetches only from GitHub-family hosts. Every GitHub row in the ledger was read in full. Cloudflare's incident reports and engineering posts were retrieved through in-session search, which returns extracted text and the canonical URL rather than the whole page, and each load-bearing number from them is cross-checked here against an independent host. The ledger in sources.md marks every row as direct or search so a reader can discount accordingly. Treat the GitHub evidence as first-hand and the rest as faithfully quoted but once removed.

02

How it is actually built

One edge machine, reconstructed from the engineering accounts, and the three separate ways something new arrives on it.

Reconstructing the request path from the public accounts gives a stack that is unusually shallow for its scale. A packet arrives at the nearest city by anycast. Inside the data centre it meets Unimog, a layer 4 load balancer running as an eBPF program on XDP, which picks a server from a forwarding table its own control plane generates; Cloudflare's eBPF account lists "Load balancer, running on XDP" as one layer of a stack of BPF programs on every edge server. Unimog exists because of a property worth dwelling on: "within a single data center, any of the servers can handle a connection for any of our services on any of our anycast IP addresses". There is no service-to-machine assignment to protect you.

On the chosen server, TLS terminates and the request enters the component Cloudflare calls the core proxy, FL. This is where each customer's configuration is applied: firewall rules, bot management, caching behaviour, routing to Workers or to R2. FL was built on NGINX, OpenResty and LuaJIT. Its replacement, FL2, is Rust on Oxy, Cloudflare's internal proxy framework, and InfoQ's report of the migration describes the sequencing choice precisely: rather than maintain two copies of the product logic, the team "added a layer inside the old NGINX/OpenResty FL that lets the new FL2 modules run". The stated motivation was not raw speed. Engineers were spending growing time working around LuaJIT bugs, and the unstructured Lua made it unsafe to add product logic.

Upstream of the proxy sits Pingora, which Cloudflare describes as the proxy that connects it to the Internet, and which is the one component of this stack you can read. Its README claims "more than 40 million Internet requests per second for more than a few years". Compute is Workers: a V8 isolate per tenant inside a shared process, not a container, because the 2018 account states the requirement as running untrusted code "securely, with low overhead" and notes that one process can hold "hundreds or thousands of Isolates". HTTP/3 at the edge comes from quiche, a separate Rust library whose README says it "powers Cloudflare edge network's HTTP/3 support".

Do not read the repositories as the system

Three pieces of direct evidence say the open-source artefacts are not the production ones. The workerd README states it "does not contain suitable defense-in-depth against the possibility of implementation bugs" and that the hosting service "uses many additional layers of defense-in-depth". Pingora's HTTP/3 support has been an open issue since March 2024 while the edge has served HTTP/3 via quiche for years. And when a contributor sent a large TLS refactor as pull request #336, the maintainer called it "a giant pr", asked for it to be marked do-not-merge, said he was recreating the changes internally, and closed it two months later with "This has been merged!". The public repository is downstream of an internal one. Architects citing Pingora in a design review are citing a library Cloudflare publishes, not the proxy that serves their traffic.

Now the part that matters, and that no single source draws. Three different kinds of change reach that machine, through three different pipelines, with three very different amounts of safety machinery around them.

Figure 2 · Three paths onto one edge server

Data path: on a timer

Config path: global in seconds

Code path: staged

Build and test

Staged release

Dashboard and API

Quicksilver

Analytics cluster
query

Feature and rule files

Core proxy FL / FL2
on every server

Pingora to origin

Workers isolate

Data path: on a timer

Config path: global in seconds

Code path: staged

Build and test

Staged release

Dashboard and API

Quicksilver

Analytics cluster
query

Feature and rule files

Core proxy FL / FL2
on every server

Pingora to origin

Workers isolate

Only the leftmost path has a documented release protocol; the configuration path reaches every machine in seconds and the data path regenerates on a timer. The proxy reads all three the same way at request time, and the incidents in section 4 arrive down the middle and right paths. Reconstructed from Cloudflare's configuration distribution, proxy and incident accounts.
Diagram source

The code path has a protocol, and you can read it. Pingora's own documentation describes a graceful upgrade in which the new instance starts with --upgrade, takes over the listening socket from the old one on SIGQUIT, and guarantees that "no request will see connection refused when trying to connect to the server endpoints". That is a designed hand-over, built because releasing code is understood to be dangerous.

The configuration path has Quicksilver, a replicated key-value store written to replace Kyoto Tycoon from 2015 onward, which fans customer configuration out to every edge machine. Its published performance characteristic is worth noting for a different reason than speed: read latency degrades sharply under concurrent writes, from a p99 of about 9 ms with no writers to about 154 ms with one writer pushing 40 kB values and about 701 ms with two. Configuration distribution is not a background concern; it competes with request serving for the same disks. The 2026 successor reports a write replication p99 of 256 ms to over 300 locations, against 4.38 s for the older mode.

The third path is the one with no name in the sources, so I will name it: the data deploy. A query runs against an analytics cluster, its output becomes a file, the file is regenerated every few minutes and shipped to every proxy, and the proxy parses it at load and uses it per request. Cloudflare's bot-management feature file is exactly this, regenerated every five minutes. Nobody calls it a deploy. It has no version the operator names, no staged rollout, no canary, and in November 2025 no schema check. It is, in every operational sense, a global code change: it alters what the proxy does to every request on earth, and it does so faster than any code release could.

The uniform fleet

Any server serves any service on any address. This is what makes the architecture cheap to operate and what removes every natural bulkhead. A bad global value has no partition to stop it.

Stated at: Cloudflare, Unimog, 2020

The core proxy as single chokepoint

FL applies every customer's security and routing config in one process. In November 2025 bypassing it was the mitigation for Workers KV and Access, which tells you how much sits behind it.

Stated at: Cloudflare, 18 November 2025

Config distribution as a tier-zero service

Quicksilver is in the request path's dependency set, not beside it. Its write load shows up as read latency on the machines serving traffic, which is why batching and compression were part of its original design.

Stated at: Cloudflare, Quicksilver, 2020

03

The decisions that matter

Eight forks taken over ten years, each with the stated reason and the condition that would flip it for you.

Decision: build your own proxy, fork the incumbent, or adopt Envoy?

Chosen
  • Write Pingora in Rust, with a bespoke HTTP library rather than hyper
  • The binding constraint was NGINX's process model: "each worker process had its own connection pool", so a request "could only reuse connections held by the worker that received it", producing redundant TLS handshakes at their volume
  • Reported outcome, Cloudflare's own measurement: roughly 70% less CPU and 67% less memory for the same traffic
Rejected
  • Forking NGINX, because extensibility was the second complaint and a fork inherits it
  • An off-the-shelf HTTP library, because their traffic includes "bizarre and non-RFC compliant HTTP traffic" they must keep serving
Flips when
  • Your connection reuse rate is already high, or handshakes are not a measurable share of CPU: then the rewrite buys nothing you can bill for
  • You need an ecosystem of ready-made filters more than you need control. Note that pingora issue #937 is the community asking for Proxy-WASM filters back, unanswered since July 2026

Decision: how fast should a configuration change reach the whole fleet?

Chosen
  • Seconds, globally, by design. Quicksilver in 2015, sub-second replication by 2026
  • The stated justification is threat response: in July 2019 the WAF rules "were deployed globally in one go, despite its normal progressive deployment procedure", so that newly disclosed vulnerabilities could be blocked quickly
Rejected
  • Routing every configuration change through the progressive rollout that code uses, on the grounds that it is too slow for security work
Flips when
  • The change is not a security emergency. The December 2025 incident was a change disabling an internal test tool, pushed through "Cloudflare's global configuration system" and propagated "within seconds". The emergency channel had become the ordinary channel, which is the failure mode to watch for in your own flag system
DecisionChosenRejectedBecauseFlips whenEvidence
Multi-tenant compute unitV8 isolate per tenant in a shared processContainer or VM per tenantStart-up cost and per-tenant memory; one process holds hundreds to thousands of isolatesThe workload needs OS APIs, native code or long-lived processes. Cloudflare's own follow-up concedes the isolate tooling ecosystem "is still young"Cloudflare, 2018
Core proxy languageRust on Oxy (FL2)Keep NGINX, OpenResty and LuaJIT (FL1)Time lost to LuaJIT bugs; unstructured Lua made product changes unsafeNever purely on language grounds: on 18 Nov 2025 the Rust path 5xx'd where Lua degraded. Rewrite only with input validation moved to ingestionInfoQ, 2025
Migration shape for the proxy rewriteRun new Rust modules inside the old proxyTwo parallel proxies with duplicated product logicAvoid maintaining two copies of every featureYour old runtime cannot host the new one. The cost of this choice is a long period where both failure modes are live, which both 2025 incidents demonstrateInfoQ, 2025
Config replication topologyFull copy on every server (Quicksilver v1)Partitioned or on-demand fetchEdge read latency; a cache miss on config is a request-path stallThe dataset outgrows the cheapest disk. Quicksilver v2 exists because "storing a full copy on every server was becoming expensive"Cloudflare, 2026
Config correctness methodDeclarative policies in a verified DSL, for DNS onlyImperative generators that build query-to-IP mapsImperative assignment hides behaviour "particularly when objectives conflict"; policies let conflicts be caught before deploymentNothing in the record flips it. The open question is why the method stayed inside DNS and never reached the rule and feature files that caused the outagesTopaz, SIGCOMM 2024
Control-plane placementPrimary facility plus disaster recovery (pre-2023)Multi-region active-activeThe control plane was not on the data plane's critical path, so its availability target was lowerWhen customers cannot change configuration or read logs during an incident. Flipped by November 2023, after which failover was the explicit fixCloudflare, 2023
Storage backend for Workers KVSingle backend, partly third-partyRedundant backends across providersShipping speed; the dependency was not treated as tier zeroWhen a downstream product makes KV a hard dependency. Flipped after 12 June 2025; Cloudflare then re-architected for redundancySquiz incident record, 2025
Object storage egress pricingZero egress (R2)Per-GB egress like the incumbentsEgress is the line item that locks customers in; removing it is a structural, not tactical, differenceStorage-heavy, read-light workloads: a third-party model found S3 73% cheaper in one of two scenarios, because R2's per-operation charges are higherVantage, 2026

Figure 3 · What gates should a change pass?

No

Yes

Yes

No

No

Yes

Does the change alter
fleet behaviour at
request time?

Normal review

Is it generated
by a machine?

Schema, size and row-count
check at publish time;
refuse to publish

Is it a security
emergency?

Stage it: rings, then
automatic halt on
error-rate delta

Push fast, but only to a
consumer that fails open
and keeps last-known-good

No

Yes

Yes

No

No

Yes

Does the change alter
fleet behaviour at
request time?

Normal review

Is it generated
by a machine?

Schema, size and row-count
check at publish time;
refuse to publish

Is it a security
emergency?

Stage it: rings, then
automatic halt on
error-rate delta

Push fast, but only to a
consumer that fails open
and keeps last-known-good

The decision tree Cloudflare's own remediation lists converge on, drawn as a rule you can apply to your own flag and ruleset pipelines. Every terminal node is an action, and the leftmost branch is the one most organisations do not build.
Diagram source
04

What broke in production

Eight published incidents, grouped into four classes. Six of the eight began with a change to configuration or machine-generated data.

Grouping these by mechanism rather than by date produces four classes, and the classes are more useful than the stories. Class A is a global configuration push with no stage gate. Class B is a latent configuration error activated by a later, unrelated change. Class C is a network configuration change whose effect depended on statement order or topology. Class D is a dependency the data plane was assumed not to have. Note what is absent: no incident in this corpus was caused by a code release that passed review and then failed. The release engineering around code works.

Figure 4 · The 18 November 2025 failure path

Customer requestCore proxy, FL2Global distributionFeature file builderAnalytics clusterCustomer requestCore proxy, FL2Global distributionFeature file builderAnalytics cluster11:05 UTCpermissions change11:28 first errors, oscillating asgoodand bad files alternate14:24 propagation halted14:30 last-known-good globalquery now reads two schemasduplicate rows, ~60 to200+ featurespublish, no schema or sizecheckevery 5 min, staggered refreshpast preallocated 200limitunwrap() panics, 5xx
Customer requestCore proxy, FL2Global distributionFeature file builderAnalytics clusterCustomer requestCore proxy, FL2Global distributionFeature file builderAnalytics cluster11:05 UTCpermissions change11:28 first errors, oscillating asgoodand bad files alternate14:24 propagation halted14:30 last-known-good globalquery now reads two schemasduplicate rows, ~60 to200+ featurespublish, no schema or sizecheckevery 5 min, staggered refreshpast preallocated 200limitunwrap() panics, 5xx
Notice that nothing between the permissions change and the panic is a deploy, a review or a gate: the only human decision in the chain happened on a database, three systems upstream of the proxy that crashed. Times from Cloudflare's published timeline.
Diagram source

Class A: the global push with no gate

Postmortem

A regular expression at 100% CPU worldwide, 2 July 2019

AssumptionA WAF rule is configuration, so it can skip the progressive deployment that code uses, because attacks do not wait.
What happenedOne misconfigured rule in a routine managed-rules deployment contained a regex that backtracked catastrophically. The rules were "deployed globally in one go, despite its normal progressive deployment procedure", and CPU hit 100% on machines worldwide.
Blast radiusRoughly 27 minutes; customer traffic down by as much as 82% at the worst point. Global WAF termination executed at 14:07, traffic back by 14:09.
FixReinstated CPU-usage protection that had been removed, added performance profiling of rules to the test suite, moved toward regex engines with run-time guarantees, and paused WAF release work.
Design ruleA rule file is a program. If a value you ship can consume unbounded CPU in your request path, the consumer needs a resource bound regardless of what the publisher promises.
Postmortem

A bot-management feature file panics the proxy, 18 November 2025

AssumptionA file your own systems generate is trustworthy, so the proxy can preallocate for its expected size and parse it without defence.
What happenedA database permissions change at 11:05 UTC made the generating query read two schemas instead of one, producing duplicate rows. The file grew from about 60 features to more than 200, past a 200-feature limit the proxy had sized memory against. A Rust function then called unwrap() on a failing result and panicked.
Blast radiusFirst errors around 11:28 UTC, propagation halted 14:24, good file global 14:30, full restoration 17:06. Availability fluctuated because the file regenerates every five minutes and proxies refreshed on staggered schedules, so good and bad files alternated.
FixNamed remediation includes validating and linting internally generated configuration before rollout, more global kill switches, stopping error reporters from exhausting resources, reviewing failure modes per proxy module, and auditing every memory preallocation and file-size limit.
Design ruleTreat output from your own pipeline as untrusted input. The schema check belongs at publish time, where refusing costs you a stale file, not at parse time on 300 sites, where refusing costs you the request.
SourceCloudflare postmortem, 18 November 2025; independently measured by ThousandEyes
Postmortem

Disabling a test tool breaks 28% of traffic, 5 December 2025

AssumptionSeventeen days after the previous incident, that a small internal flag change was not the kind of change the new safeguards were for.
What happenedWhile raising the WAF request-body buffer from 128 KB to 1 MB to mitigate a React Server Components vulnerability, an internal WAF testing tool could not handle the larger buffer, so engineers disabled it "through Cloudflare's global configuration system". The change "propagated across the fleet within seconds" and hit a long-standing bug in the Lua rules module, which tried to evaluate a non-existent action and returned 500s.
Blast radius08:47 to 09:12 UTC, about 25 minutes, affecting roughly 28% of all HTTP traffic Cloudflare serves: only customers on FL1 running the managed ruleset.
FixCloudflare states the impact "could have been less severe if the planned safeguards had already been fully implemented", naming stricter versioning and rollback, break-glass procedures and fail-open handling for configuration errors, and says it suspended network changes until those systems are complete.
Design ruleA latent bug in a config consumer is a loaded gun that your fast-propagation path fires. After an incident of this class, freeze the channel, not just the change: the second incident came through the same channel while the fixes were still in flight.
Source

The same failure class, open in the public repository

AssumptionThat unwrap() in a long-running proxy is a code-style question rather than an availability one.
What happenedIn February 2026, pingora issue #805 argued that "in a production-grade proxy/load-balancer like Pingora, relying on unwrap() is risky as it can lead to unrecoverable panics", naming peer.rs and health_check.rs and proposing Clippy's unwrap_used lint. It was closed with no visible closing comment. In June 2026 issue #921 reopened the same ground for HttpPeer and HttpHealthCheck, where an unreachable DNS server panics a worker. In September 2026 issue #1023 found one in the logging path, on non-UTF-8 bytes.
Blast radiusNo incident attributed to these in public. The point is the pattern: three independent reports of one failure class in ten months, two still open.
FixProposed, not landed: return Result, propagate with ?, and lint against new unwrap() calls. PR #920 is open.
Design ruleBan panicking constructors in anything that parses external or generated input, and enforce it with a lint rather than review. The November 2025 outage is what this class looks like when the input is a file instead of a DNS answer.
Sourcepingora #805, #921, #1023, read 2026-10-08

Classes B, C and D: delayed activation, ordering, and the dependency nobody counted

Postmortem

1.1.1.1 withdrawn by a five-week-old error, 14 July 2025

AssumptionThat a configuration change with no observable effect has no effect.
What happenedA 6 June change preparing a Data Localization Suite service put the resolver's prefixes into a non-production service topology; "the network configuration was not changed at this time", so nothing happened. On 14 July, adding a test location to that non-production service triggered a global refresh, which withdrew the resolver's routes.
Blast radius62 minutes, 21:52 to 22:54 UTC. Re-advertising routes recovered only about 77% of traffic because roughly 23% of edge servers had already removed the required local IP bindings, which had to be restored separately.
FixCloudflare attributes the cause to "a misconfiguration of legacy systems used to maintain the infrastructure that advertises Cloudflare's IP addresses", and the remediation is the retirement of that mechanism.
Design ruleAny config change that is inert today is a scheduled incident with no date. Validate against the full topology at write time, not at refresh time, and alarm on a config that references production resources from a non-production object.
Postmortem

Statement order takes 19 sites offline, 21 June 2022

AssumptionThat a resilience project's rollout is itself low risk, and that the newer architecture is the safer place to be.
What happenedA BGP prefix-advertisement policy change was deployed with its configuration statements in the wrong order, which deleted a required subset of prefixes. The sites that failed were precisely those already migrated to the newer "more flexible and resilient architecture" Cloudflare calls Multi-Colo PoP.
Blast radius19 of the busiest data centres, 06:27 to 07:42 UTC. Cloudflare states plainly: "this was our error and not the result of an attack".
FixStage-and-process changes around network configuration deployment, with a staged rollout across locations.
Design ruleMigrating to a more resilient architecture concentrates risk in the migrated set until the rollout tooling is as mature as the design. Order-sensitive configuration languages need a diff preview that shows the resulting state, not the submitted statements.
Postmortem

One router config draws the whole backbone, 17 July 2020

AssumptionThat a local change to relieve congestion in one city has local effect.
What happenedWhile working an unrelated problem on the Newark to Chicago segment, engineers changed a router configuration in Atlanta. "That configuration contained an error that routed all backbone traffic to Atlanta." The router saturated and the backbone-connected sites failed.
Blast radius27 minutes; traffic across the network down about 50%; 19 named locations affected while the rest ran normally. Resolved by removing the Atlanta router from the backbone.
FixA global change to backbone configuration. Matthew Prince framed the deeper gap as the absence of systems to stop a single mistake becoming widespread.
Design ruleIn a network, configuration blast radius is a routing property, not an organisational one. Any change that can alter path selection needs simulation against the current topology before commit.
Postmortem

A utility fault takes the control plane down for days, 2 November 2023

AssumptionThat because the control plane is not on the data plane's critical path, it can run with a lower availability design and an untested failover.
What happenedAn unplanned utility maintenance event at 08:50 UTC cut one of two independent feeds into a third-party colocation facility that DatacenterDynamics reports was the primary site for the control plane and analytics. UPS runtime fell short and generators did not start in time; The Register describes the site going from two utility sources to one plus generators, to batteries, to nothing. Failover to other facilities then did not work as expected.
Blast radiusControl plane and analytics out from 11:43 UTC on 2 November; most restored at the disaster-recovery site by 17:57 UTC; the timeline ends 04:25 UTC on 4 November. Raw logs unavailable to most customers throughout. Network and security services "continued to work as expected".
FixMake the control plane survive the loss of any single facility, and the Code Orange process, which reassigns most engineering to the crisis at hand.
Design ruleDuring an incident the control plane is the critical path, because it is how you change configuration and how you see what is happening. Rate it by what you need during a failure, not by what serves traffic.
Postmortem

Workers KV's storage dependency fails, 12 June 2025

AssumptionThat a key-value store used as a cache will be used as a cache, and that one storage backend is enough for a service other products hard-depend on.
What happenedStorage infrastructure beneath Workers KV failed. Squiz's incident record states Cloudflare said part of that infrastructure "is backed by a third-party cloud provider" which had its own outage that day. Access, WARP and the dashboard went with it. Logto's own postmortem shows the customer-side mechanism: its Worker cached region mappings in KV and "threw errors instead of falling back to uncached behavior".
Blast radiusAbout 2 hours 28 minutes by Cloudflare's reckoning; Squiz logged customer impact from 18:10 UTC on 12 June to 03:01 UTC on 13 June, well past Cloudflare's resolution at 20:28 UTC.
FixCloudflare re-architected KV's storage backend for redundancy, to remove the single point of failure. Logto re-added caching with a graceful fallback.
Design ruleA cache that throws on miss is not a cache, it is a dependency with a latency benefit. Test the uncached path in production regularly, and the provider's recovery time is not your recovery time.
The gap in this record, and it is the important one

No public account describes a Cloudflare configuration rollout that a stage gate caught. The record contains only the pushes that failed. So nobody outside the company can say whether the controls promised in November and December 2025 work, and the same is almost certainly true of your own flag pipeline: you have evidence of the incidents it caused and none of the incidents it prevented. The only way to produce that evidence is to inject a bad value deliberately and watch the gate hold.

05

Numbers you can plan against

Scale, propagation latency, incident duration and the one cost figure that bears on architecture. Every row carries its date.

MetricValueKindContextAs ofSource
Interconnection footprint500+ Tbps, 330+ citiesReportedSum of transit, peering, exchange and interconnect ports2026-04Cloudflare
Proxy throughput40M+ req/sReportedPingora, "for more than a few years"2026-10pingora README
Proxy rewrite efficiency~70% less CPU, ~67% less memoryClaimedPingora against the previous service at equal load; Cloudflare's own measurement2022-09Phoronix, reporting Cloudflare
Config write replication, current107ms p50, 181ms p95, 256ms p99ReportedQuicksilver v2 "Instant" mode to all edge locations, over 300 at publication2026-10Cloudflare
Config write replication, legacy mode4.38s p99ReportedSame table, "Classic" mode2026-10Cloudflare
Config read latency under write load9ms → 154ms → 701ms p99MeasuredQuicksilver v1 with zero, one and two concurrent writers at 40 kB values2020Cloudflare
Feature file size at failure~60 → 200+ features, limit 200ReportedBot management file, duplicated rows from a two-schema query2025-11-18Cloudflare
Data-deploy regeneration interval5 minReportedWhy availability oscillated: staggered refresh across the fleet2025-11-18Cloudflare
Change to first customer error~23 minDerived11:05 UTC change, errors from about 11:28 UTC; arithmetic on the published timeline2025-11-18Cloudflare
Time to full restoration~6h 1minDerived11:05 to 17:06 UTC; most services recovered at 14:302025-11-18Cloudflare
Blast radius of one flag change28% of HTTP traffic, 25 minReportedFL1 customers on the managed ruleset; 08:47 to 09:12 UTC2025-12-05Cloudflare
Worst traffic loss in the corpusup to 82% down, ~27 minReportedThe 2019 WAF regex incident2019-07-02Cloudflare
Partial recovery ceiling~77% on route restoreReported1.1.1.1; the other 23% of edge servers needed IP bindings restored2025-07-14Cloudflare
Verified-config query rate~1M DNS queries/sMeasuredTopaz executing declarative policies live at a global CDN2024-08SIGCOMM paper
Object storage egress$0/GB vs $0.09/GBClaimedR2 against S3 internet egress; R2's per-operation charges are higher and one of two modelled scenarios favours S3 by 73%2026Vantage
Read these carefully

Three cautions. The Pingora efficiency figures are Cloudflare's own measurement reported second-hand, with no independent benchmark in this corpus, so treat them as a claim. The 2026 replication latencies come from a product announcement describing the Workers KV path, which shares Quicksilver but is not identical to internal configuration distribution; the inference that internal config propagates on the same order is mine, and it is consistent with the December 2025 report of a change propagating "within seconds". The cost row is a third-party model, not either vendor's price list.

Three quantities nobody has published, which is where your risk sits if you are planning against this architecture: the share of traffic now on FL2 rather than FL1, the number of distinct machine-generated files the core proxy reads at request time, and any measurement of how long a configuration rollback takes end to end. The last one is the number that would tell you whether the November 2025 remediation is real.

06

The evidence wall

Every source behind this page, graded. GitHub rows were fetched in full in this session; the rest were retrieved through in-session search, as section 1 explains. Filter by kind.

Postmortem Cloudflare2025-11

Cloudflare outage on November 18, 2025

A data deploy as an outage mechanism, end to end: a permissions change upstream, duplicate rows, a file past a preallocated limit, a panic. The only document here that compares two proxy generations failing on the same input.

Carry forwardValidate generated config at publish time, and make the consumer fail open to last-known-good.
https://blog.cloudflare.com/18-november-2025-outage/
Postmortem Cloudflare2025-12

Cloudflare outage on December 5, 2025

The sequel, seventeen days later, and the most useful document here: Cloudflare names which safeguards were not yet in place, namely versioning and rollback, break-glass, and fail-open handling of configuration errors.

Carry forwardAfter an incident of this class, freeze the channel rather than the individual change.
https://blog.cloudflare.com/5-december-2025-outage/
Postmortem Cloudflare2019-07

Details of the Cloudflare outage on July 2, 2019

The origin of the pattern: a managed WAF rule skipped progressive deployment by design and a backtracking regex saturated CPU worldwide. The 2019 remediation list and the 2025 one are the same list.

Carry forwardBound the resources any shipped rule can consume in the request path.
https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019
Postmortem Cloudflare2025-07

Cloudflare 1.1.1.1 incident on July 14, 2025

The best published example of a latent configuration error: planted 6 June with no effect, activated five weeks later by a test location added to an unrelated non-production service. Route restoration alone did not recover service.

Carry forwardValidate config against the whole topology at write time; an inert change is a dated incident without the date.
https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025
Postmortem Cloudflare2023-11

Post mortem on the Cloudflare Control Plane and Analytics Outage

A utility event at one third-party facility removed the control plane and analytics for two days while the data plane kept serving. The honest part is the admission that failover did not work as expected.

Carry forwardRate the control plane by what you need during an incident, not by whether it serves traffic.
https://blog.cloudflare.com/post-mortem-on-cloudflare-control-plane-and-analytics-outage
Postmortem Cloudflare2022-06

Cloudflare outage on June 21, 2022

A BGP policy change with misordered statements deleted required prefixes in 19 of the busiest locations, which were exactly the sites already migrated to the newer resilient architecture.

Carry forwardA resilience migration concentrates risk in the migrated set until its rollout tooling matures.
https://blog.cloudflare.com/cloudflare-outage-on-june-21-2022
Postmortem Cloudflare2020-07

Cloudflare outage on July 17, 2020

A congestion fix in Atlanta routed all backbone traffic to Atlanta. Useful as the clearest statement that configuration blast radius in a network is set by routing, not by the operator's intent.

Carry forwardSimulate path-selection changes against live topology before commit.
https://blog.cloudflare.com/cloudflare-outage-on-july-17-2020
Postmortem Logto2025-06

Postmortem, June 12, 2025

The most instructive non-Cloudflare document here: this customer's Worker cached region mappings in Workers KV and threw errors rather than falling back when KV went away.

Carry forwardA cache that throws on miss is a dependency. Exercise the uncached path in production.
https://blog.logto.io/postmortem-june-12-2025
Postmortem Squiz2025-06

Service degradation: all services behind Cloudflare

Independent record of the Workers KV incident: Cloudflare's statement that part of the storage was backed by a third party, and a customer impact window running hours past Cloudflare's resolution time.

Carry forwardYour provider's resolution timestamp is not your recovery time; measure your own.
https://status.squiz.cloud/incidents/ly6brlc1jm78
Source cloudflare/pingora2026-02

Issue #805: replace excessive .unwrap() calls with robust error handling

Argued nine months before the November 2025 outage demonstrated the class: unwrap in a long-running proxy leads to unrecoverable panics. Proposes Result returns and Clippy's unwrap_used lint. Closed with no visible reason.

Carry forwardEnforce no-panic in input-parsing code with a lint, not with review.
https://github.com/cloudflare/pingora/issues/805
Source cloudflare/pingora2026-06

Issue #921: avoid panics by removing unwraps in HttpPeer and HttpHealthCheck

The same class, still open, with a concrete trigger: an unreachable DNS server panics a worker through unwrapped to_socket_addrs(). PR #920 open, so the argument is unresolved rather than settled.

Carry forwardConstructors that can fail must return Result, especially in health-check paths.
https://github.com/cloudflare/pingora/issues/921
Source cloudflare/pingora2026-09

Issue #1023: panic in a trace! call on non-UTF-8 bytes

Third instance of the class in ten months, and the most pointed: the panic is in the logging path, so observability code can kill the request it is describing.

Carry forwardDiagnostics must be the most defensive code in the system, not the least.
https://github.com/cloudflare/pingora/issues/1023
Decision cloudflare/pingora2024-10

PR #336, closed unmerged: rustls at compile time

A large external TLS refactor, marked do-not-merge at the maintainer's request, reimplemented internally, then closed as merged. Stated reasons: review size and limiting "how fast the library changes for production users".

Carry forwardRead an open-sourced component as downstream of an internal one unless the project says otherwise.
https://github.com/cloudflare/pingora/pull/336
Decision cloudflare/pingora2026-07

RFC #937: a Proxy-WASM dynamic filter subsystem

Six years after Cloudflare left Lua for Rust, a contributor proposes hot-loadable WebAssembly filters so operators can ship proxy logic without recompiling. Open, unanswered, and exactly the code-versus-config boundary this guide is about.

Carry forwardRuntime-loadable logic buys deploy speed and re-imports the blast radius you removed.
https://github.com/cloudflare/pingora/issues/937
Source cloudflare/pingora2026-10

pingora README and docs/user_guide/graceful.md

The scale claim (40M+ requests per second) and the code-release protocol: the new instance takes the listening socket on SIGQUIT so no request sees a refused connection. A designed hand-over for code, with no documented counterpart for config content.

Carry forwardCompare the care in your release path with the care in your config path; the asymmetry is the finding.
https://raw.githubusercontent.com/cloudflare/pingora/main/docs/user_guide/graceful.md
Source cloudflare/workerd2026-10

workerd README

States that the published runtime "does not contain suitable defense-in-depth against the possibility of implementation bugs" while the hosting service adds many more layers: a candid statement of the gap between an open artefact and its service.

Carry forwardSelf-hosting a vendor's runtime inherits its features, not its operational hardening.
https://github.com/cloudflare/workerd
Source cloudflare/quiche2026-10

quiche README

HTTP/3 at the edge is served by this library, separate from the open-source proxy. Read alongside pingora issue #95, open since March 2024, asking for HTTP/3 support.

Carry forwardProtocol support in a vendor's service tells you nothing about its open components.
https://raw.githubusercontent.com/cloudflare/quiche/master/README.md
Eng blog Cloudflare2022-09

How we built Pingora, the proxy that connects Cloudflare to the Internet

The rewrite rationale, and it is not per-request speed: NGINX's per-worker connection pools meant a request could only reuse connections held by its own worker, producing redundant handshakes at their volume.

Carry forwardFind the structural constraint before rewriting; "faster" is rarely the real reason.
https://blog.cloudflare.com/how-we-built-pingora-the-proxy-that-connects-cloudflare-to-the-internet/
Eng blog Cloudflare2020-04

Introducing Quicksilver: configuration distribution at Internet scale

Why the previous store was replaced from 2015, plus the measurement that matters most here: config read p99 degrades from about 9 ms to about 701 ms with two concurrent writers, so distribution competes with serving.

Carry forwardConfig distribution is a tier-zero dependency with its own saturation curve, not background work.
https://blog.cloudflare.com/introducing-quicksilver-configuration-distribution-at-internet-scale/
Eng blog Cloudflare2020-09

Unimog, Cloudflare's edge load balancer

The uniformity statement that explains the blast radius of everything else: inside a data centre any server can handle any service on any anycast address. An eBPF and XDP layer 4 balancer whose control plane generates the forwarding tables.

Carry forwardFleet uniformity is an efficiency win that deletes your natural bulkheads; add them deliberately.
https://blog.cloudflare.com/unimog-cloudflares-edge-load-balancer/
Eng blog Cloudflare2018-11

Cloud Computing without Containers, and Containers on the Edge

The isolate decision and, in the companion post, its honest cost: the tooling ecosystem for isolates was young. Together, the clearest statement of when the choice flips toward containers.

Carry forwardPick the isolation unit by start-up cost and per-tenant memory, then budget for the ecosystem you give up.
https://blog.cloudflare.com/cloud-computing-without-containers/
Eng blog Cloudflare2023-03

Oxy is Cloudflare's Rust-based next generation proxy framework

Names the production framework FL2 is built on, which is how you know Pingora is not the thing applying your WAF rules.

Carry forwardIdentify which of a vendor's several proxies is actually in your path before reasoning about its failure modes.
https://blog.cloudflare.com/introducing-oxy/
Eng blog Cloudflare2024-11

How we prevent conflicts in authoritative DNS configuration using formal verification

The company does validate configuration formally, with a custom Lisp-like language and a verifier built on Racket and Rosette, for DNS. Nothing in the record extends the method to rule files or feature files.

Carry forwardConfig verification applied to one subsystem is a capability, not a policy; name the other subsystems.
https://blog.cloudflare.com/tag/formal-methods
Eng blog InfoQ2025-10

Cloudflare's Rust rewrite of its core proxy

The migration shape: new Rust modules hosted inside the old NGINX and OpenResty proxy, to avoid two copies of product logic. Motivation given as time lost to LuaJIT bugs and unsafe Lua.

Carry forwardHosting the new runtime inside the old one avoids duplication and keeps both failure modes live for years.
https://www.infoq.com/news/2025/10/cloudflare-rust-proxy
Eng blog Gavin Howard2025-12

Piecemeal formal verification: Cloudflare, Java exceptions and Rust mutexes

The dissenting read: large organisations will not adopt formal methods wholesale, but should on their most critical paths. Included because it disagrees with the conclusion that better process alone is the answer.

Carry forwardPick the two or three paths where a wrong value is catastrophic and verify those, not everything.
https://gavinhoward.com/2025/12/piecemeal-formal-verification-cloudflare-java-exceptions-and-rust-mutexes
Paper Cloudflare Research2024-08

Topaz: Declarative and Verifiable Authoritative DNS at CDN-Scale (SIGCOMM)

Imperatively generated query-to-IP maps hide nameserver behaviour when objectives conflict; encoding objectives as policies in a formally verified DSL catches conflicts before deployment. Runs at roughly one million queries per second.

Carry forwardWhere config is generated by a program, verify the program's intent, not just the output's syntax.
https://research.cloudflare.com/publications/Larisch2024
Case study ThousandEyes2025-11

Cloudflare Outage Analysis: November 18, 2025

The only independent view in this corpus of how a global config failure looks from outside: not regional, and not a clean on-off.

Carry forwardBuy or build an outside-in measurement, because your own telemetry rides the thing that is failing.
https://www.thousandeyes.com/blog/cloudflare-outage-analysis-november-18-2025
Talk Kenton Varda, QCon and InfoQ2018-2019

Fine-grained sandboxing with V8 isolates

The isolate architecture defended by its author: multi-tenancy without VMs or containers, with claimed cold starts 10 to 100 times faster. No timestamp, because the recording could not be opened in this session.

Carry forwardThe isolate bet is a bet on per-tenant memory; model that number before choosing.
https://www.infoq.com/presentations/cloudflare-v8
Talk Cloudflare TV2020-2021

Quicksilver: Configuration Distribution at Internet Scale

A session on the configuration distribution system itself, listed because the mechanism is central to this guide's argument. The recording could not be opened in this session, so no claim rests on it.

Carry forwardWhen a vendor gives a talk about its config plane, that plane is load-bearing enough to ask about.
https://cloudflare.tv/event/3U5jrbEysk9yNFCPDwuCik
07

Build a miniature, then productionise it

Six rungs that turn this reading into a tested control in your own config pipeline. The line from toy to real is between rungs three and four.

Ship a value to a fleet and time it

Three containers reading one JSON file from a shared store on a two-second poll, each serving a request that consults it. Change the file and record when each replica's behaviour changes.

Done when: you can state your own propagation p99 from write to last replica.  Teaches: you now own a global-change mechanism with no gate on it, which is the starting position of every incident in section 4.

Break it the way Cloudflare broke it

Publish a file with a duplicated section so it exceeds whatever your consumer preallocated, and a second one with a field of the wrong type. Do not add handling yet. Watch what the replicas do.

Done when: you have reproduced both a crash and a wrong-answer failure from the same channel.  Teaches: the same bad value produces different failures in different consumers, which is why the November 2025 report shows FL2 returning 5xx while FL1 returned zeros.

Validate at publish, not at parse

Put a schema, a size ceiling and a row-count sanity bound in the publisher. Refuse to publish a file that fails any of them, and emit a metric when you refuse. Keep the previous file in place.

Done when: the bad files from rung two are rejected before distribution and the fleet keeps serving the last good one.  Teaches: the cheapest place to stop a bad value is the one place it exists once.

Make the consumer fail open and versioned

Give every published file a content hash the operator can name. On parse failure the consumer logs, increments a counter, and keeps the last known good version rather than erroring. Add a one-command rollback to any previous hash, and time it.

Done when: you can state your rollback time from decision to last replica, and a deliberately corrupt file causes zero failed requests.  Teaches: the three controls Cloudflare named after December 2025, versioning, rollback and fail-open, are one mechanism and none of them works alone.

Stage the data deploy

Add rings: one replica, then ten percent, then the rest, with an automatic halt if the error rate or a behavioural metric in the current ring diverges from the previous one. Include a documented break-glass path that skips the rings, and log every use of it.

Done when: a bad value injected at ring one never reaches ring two, and the halt is visible in your dashboard.  Teaches: why the emergency channel becomes the normal channel unless using it is recorded and reviewed.

Prove it on a schedule

Make rung two a game day you run monthly against the real pipeline, injecting an oversized file and a type error. Track time to halt and whether anyone had to intervene.

Done when: two consecutive exercises halt automatically with no human action, and you have a trend line.  Teaches: the evidence nobody in section 4 has: that your gate holds. An untested gate and no gate have the same expected value.

Figure 5 · The states a config consumer should have

last known good loaded

new version offered

schema, size and
row count pass

any check fails

operator rolls back
or a good version arrives

no validation,
no fallback state

requests fail

Serving

Validating

Serving last known good,
alarm raised, metric incremented

Crashed

last known good loaded

new version offered

schema, size and
row count pass

any check fails

operator rolls back
or a good version arrives

no validation,
no fallback state

requests fail

Serving

Validating

Serving last known good,
alarm raised, metric incremented

Crashed

The state that matters is Degraded: serving the last good value while loudly failing the load. Both 2025 incidents show consumers that had no such state, one going straight to 5xx and one to a wrong answer. Drawn from the remediation items Cloudflare named after 18 November and 5 December 2025.
Diagram source

The transferable conclusion is narrower than "test your configuration". It is this: in a uniform fleet, the fastest path to global behaviour change is the one that carries data, and it is the path least likely to have a release protocol, because nobody involved thinks of it as a release. Cloudflare spent ten years making its data plane better and its configuration plane faster, and the public record shows the second investment repaying in capability and billing in outages. If you cannot slow the propagation, and for good reasons you often cannot, then every remaining control lives at the content boundary: validate where the value is produced, version it so a human can name and revoke one, and make the consumer prefer a stale answer to no answer.

08

Keep hunting

The queries that actually produced the material above. The company-specific ones generalise by swapping the name; the vocabulary ones are the valuable half.

Incident reports with a named mechanism

  • <company> outage <date> postmortem root cause configuration file
  • <company> "incident report" "propagated" OR "rolled out globally" seconds
  • <company> outage "this was our error and not the result of an attack"
  • <company> postmortem "what we are doing" remediation "kill switch"

The argument inside the code

  • repo:<org>/<proxy> is:issue panic unwrap production stability
  • repo:<org>/<repo> is:pr is:closed is:unmerged sort:comments-desc
  • repo:<org>/<repo> is:issue "[RFC]" OR "[DNM]"
  • <repo> README "does not contain" defense-in-depth production

Architecture evolution, in the company's own words

  • <company> engineering "we replaced" OR "we outgrew" proxy CPU memory
  • <company> "configuration distribution" OR "config propagation" latency p99
  • <company> blog "began to show its age" OR "was becoming expensive"
  • <company> research publications SIGCOMM OR NSDI verifiable declarative configuration

The outside view

  • status.<customer>.com incident "behind Cloudflare" OR "upstream provider"
  • <incident date> outage analysis vantage points measured independently
  • <company> outage customer postmortem "we removed the caching" fallback
09

References

  1. Cloudflare, Details of the Cloudflare outage on July 2, 2019 Cloudflare blog, 12 July 2019. Checked 2026-10-08.
  2. Cloudflare, Cloudflare outage on July 17, 2020 Cloudflare blog, 18 July 2020. Checked 2026-10-08.
  3. Gigazine, contemporaneous report of the July 2020 outage timeline Gigazine, 20 July 2020. Checked 2026-10-08.
  4. Cloudflare, Cloudflare outage on June 21, 2022 Cloudflare blog, 21 June 2022. Checked 2026-10-08.
  5. Cloudflare, Post mortem on the Cloudflare Control Plane and Analytics Outage Cloudflare blog, November 2023. Checked 2026-10-08.
  6. The Register, Cloudflare datacenter outage The Register, 7 November 2023. Checked 2026-10-08.
  7. DatacenterDynamics, Cloudflare claims Flexential data center outage was behind service disruption DatacenterDynamics, November 2023. Checked 2026-10-08.
  8. Cloudflare, Major data center power failure (again): Cloudflare Code Orange tested Cloudflare blog, 2024. Checked 2026-10-08.
  9. Squiz, Service degradation: all services behind Cloudflare Squiz status page, 12 June 2025. Checked 2026-10-08.
  10. Logto, Postmortem, June 12, 2025 Logto blog, June 2025. Checked 2026-10-08.
  11. Cloudflare, Cloudflare 1.1.1.1 incident on July 14, 2025 Cloudflare blog, July 2025. Checked 2026-10-08.
  12. Cloudflare, Cloudflare outage on November 18, 2025 Cloudflare blog, 18 November 2025. Checked 2026-10-08.
  13. ThousandEyes, Cloudflare Outage Analysis: November 18, 2025 ThousandEyes blog, November 2025. Checked 2026-10-08.
  14. Cloudflare, Cloudflare outage on December 5, 2025 Cloudflare blog, 5 December 2025. Checked 2026-10-08.
  15. Gavin Howard, Piecemeal formal verification: Cloudflare, Java exceptions and Rust mutexes gavinhoward.com, December 2025. Checked 2026-10-08.
  16. Cloudflare, How we built Pingora, the proxy that connects Cloudflare to the Internet Cloudflare blog, September 2022. Checked 2026-10-08.
  17. Phoronix, Cloudflare ditches Nginx for in-house, Rust-written Pingora Phoronix, September 2022. Checked 2026-10-08.
  18. InfoQ, Cloudflare's Rust rewrite of its core proxy InfoQ, October 2025. Checked 2026-10-08.
  19. Cloudflare, Oxy is Cloudflare's Rust-based next generation proxy framework Cloudflare blog, 2 March 2023. Checked 2026-10-08.
  20. Cloudflare, Introducing Quicksilver: Configuration Distribution at Internet Scale Cloudflare blog, 2020. Checked 2026-10-08.
  21. Cloudflare, Quicksilver v2: evolution of a globally distributed key-value store (Part 1) Cloudflare blog, 2026. Checked 2026-10-08.
  22. Cloudflare, Introducing Workers KV Instant, powered by Quicksilver Cloudflare blog, October 2026. Checked 2026-10-08.
  23. Cloudflare, Unimog, Cloudflare's edge load balancer Cloudflare blog, 9 September 2020. Checked 2026-10-08.
  24. Cloudflare, Cloudflare architecture and how BPF eats the world Cloudflare blog, 2019. Checked 2026-10-08.
  25. Cloudflare, Cloud Computing without Containers Cloudflare blog, 9 November 2018. Checked 2026-10-08.
  26. Cloudflare, Containers on the Edge Cloudflare blog, 2018. Checked 2026-10-08.
  27. Cloudflare, Crossing 500 Tbps of capacity Cloudflare blog, April 2026. Checked 2026-10-08.
  28. Cloudflare, How we prevent conflicts in authoritative DNS configuration using formal verification Cloudflare blog, formal methods tag, November 2024. Checked 2026-10-08.
  29. Larisch, Alberdingk Thijm, Ahmad, Wu, Arnfeld, Fayed, Topaz: Declarative and Verifiable Authoritative DNS at CDN-Scale ACM SIGCOMM, August 2024. Checked 2026-10-08.
  30. cloudflare/pingora, repository and README GitHub, read at HEAD 2026-10-08.
  31. cloudflare/pingora, docs/user_guide/graceful.md GitHub, read at HEAD 2026-10-08.
  32. cloudflare/pingora issue #805, Replace excessive .unwrap() calls with robust error handling for production stability GitHub, opened 4 February 2026, closed. Checked 2026-10-08.
  33. cloudflare/pingora issue #921, Avoid panics by removing unwraps during HttpPeer and HttpHealthCheck initialization GitHub, opened 20 June 2026, open. Checked 2026-10-08.
  34. cloudflare/pingora issue #1023, Panic in trace! call due to unwrap() on non-UTF-8 bytes GitHub, opened 29 September 2026, open. Checked 2026-10-08.
  35. cloudflare/pingora pull request #336, [DNM] Add Rustls compile time implementation GitHub, closed unmerged 28 October 2024. Checked 2026-10-08.
  36. cloudflare/pingora issue #937, [RFC] Add WebAssembly (Proxy-WASM) Dynamic Filter Subsystem GitHub, opened 22 July 2026, open. Checked 2026-10-08.
  37. cloudflare/pingora issue #95, HTTP3/QUIC Support GitHub, opened 2 March 2024, open. Checked 2026-10-08.
  38. cloudflare/quiche, README GitHub, read at HEAD 2026-10-08.
  39. cloudflare/workerd, repository and README GitHub, read at HEAD 2026-10-08.
  40. Kenton Varda, Fine-grained sandboxing with V8 isolates QCon, hosted by InfoQ, 2018 to 2019. Checked 2026-10-08.
  41. Cloudflare TV, Quicksilver: Configuration Distribution at Internet Scale Cloudflare TV, 2020 to 2021. Checked 2026-10-08.
  42. Vantage, Cloudflare R2 vs AWS S3 Vantage blog, 2026. Checked 2026-10-08.