Evidence ledger
One row per claim in Hardening the wrong plane: ten years of Cloudflare's architecture: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.
Field guide: Hardening the wrong plane — ten years of Cloudflare's architecture, 2016 to 2026 Research date: 2026-10-08. Evidence cutoff: October 2026.
Retrieval method, and its limit
This matters for how much weight to put on each row, so it goes first.
This session ran behind an egress policy that permits only GitHub-family hosts
(github.com, raw.githubusercontent.com) and package registries. Every other host, including
blog.cloudflare.com, answered 403 to CONNECT at the proxy. Two consequences:
- Rows marked
fetched: directwere retrieved in full in this session and the quoted text is copied from the page. These are the GitHub rows: repository READMEs, in-repo documentation, issue and pull-request threads. - Rows marked
fetched: searchwere retrieved through in-session web search, which returns extracted content and the canonical URL but not the whole page. The quote column for these rows carries the text the search tool returned from the source; where it returned a paraphrase rather than a sentence, the quote column saysparaphraseand the claim is weakened accordingly in the guide. No URL in this ledger was written from memory: each one appeared in a search result index in this session, and eight of the Cloudflare incident URLs also appear in earlier digs in this repository where they were link-checked directly. - Where a
searchrow carries a load-bearing number, the guide cross-checks it against an independent host (ThousandEyes, The Register, DataCenterDynamics, a customer status page) and says so in the prose.
verify.mjs will report every non-GitHub link as unreachable for the same reason. Those warnings
are accepted, and they are a property of this session's network, not of the sources.
A second, different limit: this is a single-company dig, so the corpus is Cloudflare-heavy by construction. Independent accounts (rows 20 to 25) exist to corroborate, not to broaden.
| # | Org | Title | Tier | Published | Checked | Fetched | URL | Claim I take from it | Supporting quote or figure |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Cloudflare | Details of the Cloudflare outage on July 2, 2019 | postmortem | 2019-07-12 | 2026-10-08 | search | https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019 | A single WAF rule, pushed globally outside the normal staged rollout, exhausted CPU across the whole network. | "The cause of this outage was deployment of a single misconfigured rule within the Cloudflare Web Application Firewall (WAF) during a routine deployment of new Cloudflare WAF Managed rules"; it "caused CPU to spike to 100% on our machines worldwide". |
| 2 | Cloudflare | Details of the Cloudflare outage on July 2, 2019 | postmortem | 2019-07-12 | 2026-10-08 | search | https://blog.cloudflare.com/details-of-the-cloudflare-outage-on-july-2-2019 | The staged rollout was deliberately bypassed so security rules could reach the edge fast. | "these WAF rules were deployed globally in one go, despite its normal progressive deployment procedure"; mitigation: "executed the global WAF termination at 14:07 and by 14:09 traffic levels and CPU were back to expected levels worldwide". |
| 3 | Cloudflare | Cloudflare outage on July 17, 2020 | postmortem | 2020-07-18 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-outage-on-july-17-2020 | A router configuration change made to relieve congestion in one city drew all backbone traffic to that city. | "That configuration contained an error that routed all backbone traffic to Atlanta"; incident "lasted 27 minutes"; "Traffic across the network dropped by about 50%". |
| 4 | Cloudflare | Cloudflare outage on June 21, 2022 | postmortem | 2022-06-21 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-outage-on-june-21-2022 | A BGP prefix-advertisement policy change, applied in the wrong statement order, took 19 of the busiest sites offline. | "affected traffic in 19 of our data centers"; outage "started at 06:27 UTC"; "by 07:42 UTC all data centers were online and working correctly"; "this was our error and not the result of an attack". |
| 5 | Cloudflare | Cloudflare outage on June 21, 2022 | postmortem | 2022-06-21 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-outage-on-june-21-2022 | The sites that failed were the ones already migrated to the newer, more resilient architecture. | The affected sites had been moved to "a more flexible and resilient architecture", internally called "Multi-Colo PoP (MCP)". |
| 6 | Cloudflare | Post mortem on the Cloudflare Control Plane and Analytics Outage | postmortem | 2023-11-04 | 2026-10-08 | search | https://blog.cloudflare.com/post-mortem-on-cloudflare-control-plane-and-analytics-outage | The control plane depended on one facility while the data plane did not; a utility fault at that facility took the control plane down for days. | "On November 2 at 08:50 UTC Portland General Electric (PGE) ... had an unplanned maintenance event affecting one of their independent power feeds into the building. That event shut down one feed into PDX-04." Control plane outage began "11:43 UTC"; most restored at disaster-recovery site "17:57 UTC"; timeline ends "November 4 at 04:25 UTC". Network and security services "continued to work as expected". |
| 7 | Cloudflare | Major data center power failure (again): Cloudflare Code Orange tested | postmortem | 2024 | 2026-10-08 | search | https://blog.cloudflare.com/major-data-center-power-failure-again-cloudflare-code-orange-tested | The organisational response to 2023 was a named process that reassigns engineering to a crisis. | Code Orange is described as a process where the company "shifts most or all engineering resources to addressing the issue at hand" during a significant event. |
| 8 | Cloudflare | Cloudflare 1.1.1.1 incident on July 14, 2025 | postmortem | 2025-07-14 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025 | A configuration error planted on 6 June stayed inert until an unrelated change activated it five weeks later. | Resolver "unavailable to the Internet starting at 21:52 UTC and ending at 22:54 UTC"; cause was "a misconfiguration of legacy systems used to maintain the infrastructure that advertises Cloudflare's IP addresses to the Internet"; the June 6 change meant "the network configuration was not changed at this time, so the routing of 1.1.1.1 was not affected". |
| 9 | Cloudflare | Cloudflare 1.1.1.1 incident on July 14, 2025 | postmortem | 2025-07-14 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-1-1-1-1-incident-on-july-14-2025 | Re-advertising routes did not fully restore service, because edge servers had already discarded local IP bindings. | Readvertisement restored traffic to "about 77 percent" of prior level; "approximately 23 percent of edge servers had already removed required IP bindings" and full recovery required restoring them. |
| 10 | Cloudflare | Cloudflare outage on November 18, 2025 | postmortem | 2025-11-18 | 2026-10-08 | search | https://blog.cloudflare.com/18-november-2025-outage/ | A database permissions change made a machine-generated feature file double in size, past a hard limit the proxy had preallocated memory for, and the proxy panicked. | Change at "11:05 UTC"; the query began "querying both the default and r0 schemas, producing a large set of duplicate feature rows"; the file normally holds "about 60 features" and "ballooned to more than 200" against a "200-feature limit"; "a Rust function hit an error and executed an unwrap() on a failing result, triggering a panic". |
| 11 | Cloudflare | Cloudflare outage on November 18, 2025 | postmortem | 2025-11-18 | 2026-10-08 | search | https://blog.cloudflare.com/18-november-2025-outage/ | The new Rust proxy failed harder than the old Lua one on the same bad input. | "Customers on FL2 saw 5xx errors for most bot-protected traffic, while FL customers saw bot scores default to zero instead." Feature file "is regenerated every five minutes"; staggered refresh produced "fluctuating availability". Timeline: propagation halted "14:24", good file deployed "14:30", full restoration "17:06". |
| 12 | Cloudflare | Cloudflare outage on November 18, 2025 | postmortem | 2025-11-18 | 2026-10-08 | search | https://blog.cloudflare.com/18-november-2025-outage/ | The remediation list is a list of release-engineering controls for data, not code. | Planned work includes treating internally generated configuration files so "they are validated & lint before rollout", expanding "global kill switches", "preventing core dumps and error reporters from exhausting resources", reviewing "failure modes across the core proxy modules", and auditing "all memory preallocation and file size limits across core systems". |
| 13 | Cloudflare | Cloudflare outage on December 5, 2025 | postmortem | 2025-12-05 | 2026-10-08 | search | https://blog.cloudflare.com/5-december-2025-outage/ | Seventeen days later, a second global configuration push hit a latent bug in the older proxy. | "began at 08:47 UTC and was resolved at 09:12"; affected "about 28% of all HTTP traffic Cloudflare serves"; the WAF body buffer was raised "from 128 KB to 1 MB"; an internal test tool was disabled "through Cloudflare's global configuration system"; the change "propagated across the fleet within seconds"; a "long-standing bug in the Lua-based rules module" made the proxy evaluate "a non-existent 'execute' action". |
| 14 | Cloudflare | Cloudflare outage on December 5, 2025 | postmortem | 2025-12-05 | 2026-10-08 | search | https://blog.cloudflare.com/5-december-2025-outage/ | Cloudflare states in its own words that the November safeguards were not yet in place, and names them. | The December 5 impact "could have been less severe if the planned safeguards had already been fully implemented"; those safeguards include "a stricter versioning and rollback system", "break glass" procedures, and "fail-open handling for configuration errors"; Cloudflare "suspended network changes until the new mitigation systems are complete". |
| 15 | Cloudflare | How we built Pingora, the proxy that connects Cloudflare to the Internet | blog | 2022-09-14 | 2026-10-08 | search | https://blog.cloudflare.com/how-we-built-pingora-the-proxy-that-connects-cloudflare-to-the-internet/ | NGINX was replaced because its process model, not its performance per request, set the CPU floor. | "each worker process had its own connection pool", so "a request could only reuse connections held by the worker that received it", producing "redundant SSL/TLS handshakes"; Pingora "employs a multi-threaded architecture"; Cloudflare chose Rust and wrote its own HTTP library rather than using "an off-the-shelf one such as hyper" because its traffic includes "bizarre and non-RFC compliant HTTP traffic". |
| 16 | Cloudflare | Introducing Quicksilver: Configuration Distribution at Internet Scale | blog | 2020-04 | 2026-10-08 | search | https://blog.cloudflare.com/introducing-quicksilver-configuration-distribution-at-internet-scale/ | Config distribution was rewritten in 2015 because the previous store did not survive network growth, and reads degrade under concurrent writes. | What "worked for the first 25 cities began to show its age as the network passed 100"; read latency p99 "about 9 ms" with no writes, "about 154 ms" with one writer adding 40 kB values, "about 701 ms" with a second writer, p99.9 "exceeded one second". |
| 17 | Cloudflare | Quicksilver v2: evolution of a globally distributed key-value store (Part 1) | blog | 2026 | 2026-10-08 | search | https://blog.cloudflare.com/quicksilver-v2-evolution-of-a-globally-distributed-key-value-store-part-1/ | Full replication of the config set to every server stopped being affordable, which is what v2 exists to change. | The redesign came about because "storing a full copy on every server was becoming expensive"; v1 had "each server hold the full dataset and update it through asynchronous replication". |
| 18 | Cloudflare | Introducing Workers KV Instant, powered by Quicksilver | blog | 2026-10-01 | 2026-10-08 | search | https://blog.cloudflare.com/workers-kv-instant/ | After three configuration-propagation outages, global propagation got faster, not slower. | Instant mode write replication: median "107 ms", p95 "181 ms", p99 "256 ms", reaching all edge locations, "over 300" at publication; Classic p99 write replication "4.38 s"; p99 read "1.62 ms" versus "287 ms". |
| 19 | Cloudflare | Unimog, Cloudflare's edge load balancer | blog | 2020-09-09 | 2026-10-08 | search | https://blog.cloudflare.com/unimog-cloudflares-edge-load-balancer/ | Inside a data centre, any server can serve any service on any anycast address, which is what makes the fleet uniform. | "Within a single data center, any of the servers can handle a connection for any of our services on any of our anycast IP addresses"; Unimog is an eBPF/XDP layer 4 balancer whose "control plane generates forwarding tables"; it supports "daisy-chaining" to keep established TCP connections during table updates. |
| 20 | Cloudflare | Cloud Computing without Containers | blog | 2018-11-09 | 2026-10-08 | search | https://blog.cloudflare.com/cloud-computing-without-containers/ | Multi-tenant compute at the edge was built on V8 isolates rather than containers or VMs, for start-up cost and per-tenant memory. | The platform "doesn't use containers or virtual machines"; the stated requirement was to "run untrusted code securely, with low overhead"; a single process can run "hundreds or thousands of Isolates". |
| 21 | Cloudflare | Containers on the Edge | blog | 2018 | 2026-10-08 | search | https://blog.cloudflare.com/containers-on-the-edge | Cloudflare's own follow-up concedes the isolate model's ecosystem cost, which is the condition that flips the decision. | The "ecosystem of tooling and technology stacks for isolates is still young and developing". |
| 22 | Cloudflare | Oxy is Cloudflare's Rust-based next generation proxy framework | blog | 2023-03-02 | 2026-10-08 | search | https://blog.cloudflare.com/introducing-oxy/ | The production proxy framework is Oxy, which is not the open-source Pingora. | Oxy is "a feature-rich proxy server tightly integrated with our internal infrastructure"; it underpins "the Zero Trust Gateway, the iCloud Private Relay second hop proxy, and the internal egress routing service". |
| 23 | Cloudflare | Cloudflare architecture and how BPF eats the world | blog | 2019 | 2026-10-08 | search | https://blog.cloudflare.com/cloudflare-architecture-and-how-bpf-eats-the-world/ | The edge server's packet path is a stack of BPF programs, with the load balancer on XDP. | The post lists "Load balancer, running on XDP" among the BPF layers on Cloudflare's edge servers. |
| 24 | Cloudflare | Crossing 500 Tbps of capacity | blog | 2026-04 | 2026-10-08 | search | https://blog.cloudflare.com/500-tbps-of-capacity/ | Current fleet scale, for sizing the blast radius of one global push. | Capacity is counted as every transit, peering, IX and CNI port "across all 330+ cities". |
| 25 | Cloudflare | How we prevent conflicts in authoritative DNS configuration using formal verification | blog | 2024-11 | 2026-10-08 | search | https://blog.cloudflare.com/tag/formal-methods | The same company does formally verify configuration, for one subsystem, with a purpose-built language and checker. | Cloudflare uses "a custom Lisp-like programming language and formal verifier (written in Racket and Rosette)" to prevent "logical contradictions in our authoritative DNS nameserver's behavior". |
| 26 | Cloudflare Research (Larisch, Alberdingk Thijm, Ahmad, Wu, Arnfeld, Fayed) | Topaz: Declarative and Verifiable Authoritative DNS at CDN-Scale | paper | 2024-08 (SIGCOMM) | 2026-10-08 | search | https://research.cloudflare.com/publications/Larisch2024 | Imperative configuration generators hide conflicting objectives; declarative policies let conflicts be caught before deployment. | The abstract says imperative assignment systems that "imperatively build query-to-IP maps" hide nameserver behaviour "particularly when objectives conflict"; because policies are "written in a formally verified domain-specific language (topaz-lang)", Topaz "can catch policy conflicts before deployment"; it handles "roughly one million DNS queries per second". |
| 28 | Cloudflare | cloudflare/pingora, README | source | repo, read at HEAD | 2026-10-08 | direct | https://github.com/cloudflare/pingora | The open-source proxy's own claim of production scale. | "serving more than 40 million Internet requests per second for more than a few years"; features list includes "Graceful reload" and "Customizable load balancing and failover strategies". |
| 29 | Cloudflare | pingora, docs/user_guide/graceful.md | source | repo, read at HEAD | 2026-10-08 | direct | https://raw.githubusercontent.com/cloudflare/pingora/main/docs/user_guide/graceful.md | Code releases get a designed hand-over of the listening socket; nothing comparable is documented for config content. | "No request will see connection refused when trying to connect to the server endpoints." The new instance starts with --upgrade, "takes over the listening socket from the old instance", and the old instance transfers it on SIGQUIT. |
| 30 | community (rxdiscovery) and Cloudflare maintainers | pingora issue #805, "Replace excessive .unwrap() calls with robust error handling for production stability" | source | 2026-02-04, closed | 2026-10-08 | direct | https://github.com/cloudflare/pingora/issues/805 | The panic-on-unwrap failure class was raised against Cloudflare's own public proxy and argued in exactly the terms the November 2025 outage then demonstrated. | "In a production-grade proxy/load-balancer like Pingora, relying on unwrap() is risky as it can lead to unrecoverable panics." Names peer.rs and health_check.rs; proposes returning Result, ?/map_err(), and enabling Clippy's unwrap_used lint. Closed with no visible closing comment. |
| 31 | community (rxdiscovery) | pingora issue #921, "Avoid panics by removing unwraps during HttpPeer and HttpHealthCheck initialization" | source | 2026-06-20, open | 2026-10-08 | direct | https://github.com/cloudflare/pingora/issues/921 | The same class is still open four months after a global outage caused by it. | "Several constructors in the codebase rely on .unwrap(), which can cause a worker process to panic unexpectedly." HttpPeer::new() unwraps to_socket_addrs() and .next(); "the panic occurs when the DNS server is unreachable". Labels: enhancement, ergonomics. Linked PR #920, open. Version 0.8.1. |
| 32 | community | pingora issue #1023, panic in trace! from unwrap() on non-UTF-8 bytes |
source | 2026-09-29, open | 2026-10-08 | direct | https://github.com/cloudflare/pingora/issues/1023 | A third instance of the same class, in the logging path, where a diagnostic can kill the request. | Title: "Panic in trace! call due to unwrap() on non-UTF-8 bytes in protocols/http/v1/client.rs:818". Open. |
| 33 | Cloudflare (johnhurt) and contributor | pingora PR #336, "[DNM] Add Rustls compile time implementation" | adr | 2024-08 to 2024-10-28, closed unmerged | 2026-10-08 | direct | https://github.com/cloudflare/pingora/pull/336 | A large external contribution was rejected on review-size grounds, reimplemented internally, and the external PR closed as "merged": the public repo is downstream of an internal one. | Maintainer called it "a giant pr", asked for a [DNM] prefix and smaller stacked PRs to "limit how fast the library changes for production users", said he was recreating the changes internally, and closed with "This has been merged! Thanks for your contributions." He "chose compile-time configuration over trait-level abstractions". |
| 34 | community | pingora issue #937, "[RFC] Add WebAssembly (Proxy-WASM) Dynamic Filter Subsystem (pingora-wasm)" | adr | 2026-07-22, open | 2026-10-08 | direct | https://github.com/cloudflare/pingora/issues/937 | Six years after Cloudflare left Lua for Rust, the community is asking for runtime-loadable filters back, and the proposal is unanswered. | Proposes "introducing an opt-in WebAssembly filter runtime crate" using wasmtime and Proxy-WASM ABI, because Envoy-style proxies "rely heavily on WebAssembly (Proxy-WASM) for dynamic security, auth, and header-mutation plugins", letting developers "write proxy filters in Rust, Go, C++, or Zig and deploy them on the fly". No maintainer response. |
| 35 | community | pingora issue #95, "HTTP3/QUIC Support" | source | 2024-03-02, open | 2026-10-08 | direct | https://github.com/cloudflare/pingora/issues/95 | The open-source proxy still lacks the protocol Cloudflare's edge has served for years, which marks the gap between the artefact and the production system. | Issue open since 2024-03-02 in the filtered issue list. |
| 36 | Cloudflare | cloudflare/quiche, README | source | repo, read at HEAD | 2026-10-08 | direct | https://raw.githubusercontent.com/cloudflare/quiche/master/README.md | HTTP/3 at the edge is served by a separate Rust library, not by the open-source proxy. | "quiche powers Cloudflare edge network's HTTP/3 support"; "Android's DNS resolver uses quiche to implement DNS over HTTP/3". |
| 37 | Cloudflare | cloudflare/workerd, README | source | repo, read at HEAD | 2026-10-08 | direct | https://github.com/cloudflare/workerd | The published Workers runtime is explicitly not the hardened production one. | workerd "is a JavaScript / Wasm server runtime based on the same code that powers Cloudflare Workers"; it "on its own does not contain suitable defense-in-depth against the possibility of implementation bugs"; the hosting service "uses many additional layers of defense-in-depth". |
| 38 | Cloudflare | cloudflare/foundations, README (read, not cited in the page) | source | repo, read at HEAD | 2026-10-08 | direct | https://raw.githubusercontent.com/cloudflare/foundations/main/README.md | Cross-service operational concerns were factored into one library, including configuration loading, which shows where the org draws the platform boundary. | "Foundations is a modular Rust library, designed to help scale programs for distributed, production-grade systems", covering logging, tracing, metrics, memory profiling, "seccomp-based syscall sandboxing", "service configuration with documentation" and a config-loading CLI helper. |
| 39 | ThousandEyes | Cloudflare Outage Analysis: November 18, 2025 | casestudy | 2025-11 | 2026-10-08 | search | https://www.thousandeyes.com/blog/cloudflare-outage-analysis-november-18-2025 | Independent external measurement of the November 2025 outage, from outside Cloudflare's own telemetry. | Independent vantage-point analysis of the same incident window; used in the guide only to corroborate that impact was global and intermittent rather than regional. (paraphrase) |
| 40 | Logto | Postmortem, June 12, 2025 | postmortem | 2025-06 | 2026-10-08 | search | https://blog.logto.io/postmortem-june-12-2025 | A customer's own postmortem shows how a dependency on Workers KV failed open or closed by accident, not by design. | Its Worker "used KV to cache region mappings, and when KV became unavailable, the Worker threw errors instead of falling back to uncached behavior"; it removed the caching logic, then re-added it "with a graceful fallback". |
| 41 | Squiz | Service degradation: all services behind Cloudflare | postmortem | 2025-06-12 | 2026-10-08 | search | https://status.squiz.cloud/incidents/ly6brlc1jm78 | Cloudflare attributed the Workers KV failure to storage infrastructure backed by a third party, and customer-visible recovery lagged Cloudflare's resolution. | Records that "Cloudflare stated that part of this infrastructure is backed by a third-party cloud provider" which had its own outage; customer impact window 18:10 on 2025-06-12 to 03:01 UTC on 2025-06-13, against Cloudflare's resolution at 20:28 UTC. |
| 42 | InfoQ | Cloudflare's Rust rewrite of its core proxy (news report on FL2) | blog | 2025-10 | 2026-10-08 | search | https://www.infoq.com/news/2025/10/cloudflare-rust-proxy | FL2 was rolled out by embedding the new Rust modules inside the old NGINX/Lua proxy rather than running two copies of the product logic. | Reports that to avoid maintaining two copies of product logic, the team "added a layer inside the old NGINX/OpenResty FL that lets the new FL2 modules run"; FL1 ran on "NGINX, OpenResty, and LuaJIT" and engineers spent growing time "working around LuaJIT bugs". |
| 43 | Phoronix | Cloudflare ditches Nginx for in-house, Rust-written Pingora | blog | 2022-09 | 2026-10-08 | search | https://www.phoronix.com/news/CloudFlare-Pingora-No-Nginx | Independent restatement of Cloudflare's efficiency claim for the proxy rewrite. | Reports that in production Pingora "consumes about 70% less CPU and 67% less memory compared to the old service with the same traffic load" (Cloudflare's measurement, reported second-hand). |
| 44 | The Register | Cloudflare datacenter outage | blog | 2023-11-07 | 2026-10-08 | search | https://www.theregister.com/2023/11/07/cloudflare_datacenter_outage/ | Independent account of the power cascade and of the failover that did not work. | Summarises the site moving "from two utility sources to one utility plus generators, then to UPS batteries, then to nothing", and that Cloudflare "found out the hard way that its failover plans from that datacenter to other facilities didn't work quite as expected". |
| 45 | DatacenterDynamics | Cloudflare claims Flexential data center outage was behind service disruption | blog | 2023-11 | 2026-10-08 | search | https://www.datacenterdynamics.com/en/news/cloudflare-claims-flexential-data-center-outage-was-behind-service-disruption/ | The facility was the primary site for the control plane and analytics, and it was a third-party colocation site. | Reports PDX-DC04 was "its primary data center for Cloudflare's control plane and analytics systems", operated by Flexential; notes other outlets name the site differently. |
| 46 | Gavin Howard | Piecemeal formal verification: Cloudflare, Java exceptions and Rust mutexes | blog | 2025-12 | 2026-10-08 | search | https://gavinhoward.com/2025/12/piecemeal-formal-verification-cloudflare-java-exceptions-and-rust-mutexes | A practitioner's argued dissent: the answer is formal methods on the most critical paths only, not everywhere. | Argues the outage "should not have happened" and that large companies are unlikely to adopt formal methods wholesale "but for their most critical systems, they should". (argument, paraphrased by the search tool) |
| 47 | Vantage | Cloudflare R2 vs AWS S3 | blog | 2026 | 2026-10-08 | search | https://vantage.sh/blog/cloudflare-r2-aws-s3-comparison | Third-party cost comparison for the zero-egress decision, including the cases where it loses. | Storage "$0.015 per GB for R2 versus starting at \(0.023 per GB for S3"; S3 internet egress "\)0.09 per GB" against R2's zero; R2 Class A operations "$0.0045 per 1,000"; and scenario-dependence: "S3 is 73% cheaper than R2" in one of two modelled scenarios. |
| 48 | InfoQ / QCon (Kenton Varda) | Fine-grained sandboxing with V8 isolates | talk | 2018-2019 | 2026-10-08 | search | https://www.infoq.com/presentations/cloudflare-v8 | The isolate decision presented and defended by its author, with the claimed start-up advantage. | Session describes "a cloud compute platform designed for massive multi-tenancy without using virtual machines or containers, but instead using V8 isolates", with "10x-100x faster cold starts and lower memory footprints". No timestamp: the recording could not be opened in this session. |
| 49 | Cloudflare TV | Quicksilver: Configuration Distribution at Internet Scale | talk | 2020-2021 | 2026-10-08 | search | https://cloudflare.tv/event/3U5jrbEysk9yNFCPDwuCik | A talk on the configuration-distribution system itself, which is the mechanism this guide argues is load-bearing. | Event exists under this title on Cloudflare's own video channel. No timestamp or transcript: the recording could not be opened in this session. |
| 50 | Cloudflare | Workers KV product changelog (read, not cited in the page) | vendor | continuous | 2026-10-08 | search | https://developers.cloudflare.com/changelog/product/kv/ | Vendor documentation of the config surfaces customers drive, for the record of what the documented rollout controls are. | Changelog index for the KV product. Used only as a pointer; no claim in the guide rests on it. (paraphrase) |
Tier mix
| Tier | Rows |
|---|---|
| postmortem | 1-14 (Cloudflare incidents), 40, 41 |
| blog | 15-25, 42-47 |
| paper | 26 |
| source | 28-32, 35-37 (38 read, not cited) |
| adr | 33, 34 |
| casestudy | 39 |
| talk | 48, 49 |
| vendor | 50 read, not cited |
Distinct hosts cited in the page: blog.cloudflare.com, research.cloudflare.com, github.com, raw.githubusercontent.com, thousandeyes.com, blog.logto.io, status.squiz.cloud, infoq.com, phoronix.com, theregister.com, datacenterdynamics.com, gigazine.net, gavinhoward.com, vantage.sh, cloudflare.tv. Fifteen.
Organisations whose engineers or operators are the authors: Cloudflare, Logto, Squiz, ThousandEyes, Vantage, plus independent technology reporting and one practitioner writing under his own name. One company is the subject; the others corroborate or dissent.
What I looked for and did not find
- No public account of a Cloudflare configuration rollout that was caught by a stage gate. The record contains the pushes that failed, never a push that was stopped. This is the single biggest hole in the evidence, and it means nobody outside Cloudflare can say whether the controls promised after November 2025 work.
- No published figure for the share of traffic on FL2 versus FL1. Searched directly; the only percentages in the record (28% in December 2025) describe impact, not migration progress.
- No Cloudflare account of applying the DNS formal-verification approach (rows 25, 26) to any other configuration domain. The transfer is the obvious next step and no source describes it.
- Only one peer-reviewed paper bears directly on this topic. Topaz (row 26) is it. Cloudflare Research publishes steadily, but on measurement, privacy and protocol work rather than on its own configuration distribution or proxy architecture, so the mechanism of the thing that keeps causing its outages has no academic treatment. A second paper that looked relevant, an arXiv item on consistent hashing with bounded loads, was dropped from this ledger because its authorship and exact title could not be confirmed in this session and the claim it would have supported was peripheral. Dropping it is the right trade: a guessed author list is worse than a thinner corpus.
- No maintainer reply on pingora #937 or #921. The arguments about dynamic filters and about removing unwraps are both unanswered in public.