Term Kind Topic What it is
Airbnb Minerva: One Definition of a Metric Minerva Metrics Platform case-study Semantic Layer Airbnb built a central metrics platform because the same business question was producing different answers depending on which dashboard you opened.
Airbnb's Service-Oriented Migration case-study Legacy Modernization Airbnb decomposed a large Rails monolith by first extracting a unified data-access layer, so that services were built on owned data rather than on shared database tables.
Airbnb: Monorail to Service-Oriented Architecture Airbnb SOA Migration case-study Application Decomposition Airbnb decomposed a large Rails monolith into services, and the difficult part was data ownership rather than code extraction.
Airflow: A Scheduler Born From Pipeline Sprawl Apache Airflow case-study Workflow Schedulers Airbnb built Airflow because cron cannot express dependencies, and a data platform's failures are mostly dependency failures.
Amazon 2002: The API Mandate Bezos Mandate case-study API & Integration A directive that all teams expose functionality only through service interfaces, with no back doors, is widely credited with making both Amazon's architecture and AWS possible.
Amazon Dynamo: Always Writeable Dynamo Paper case-study Consistency Models Amazon chose an always-writeable shopping cart with application-level conflict resolution, accepting merge complexity to guarantee that "add to cart" never fails.
Amazon Prime Video: Serverless Back to a Monolith Prime Video Audio/Video Monitoring case-study Monolith vs Microservices A distributed serverless pipeline was consolidated into a single process, reducing cost by over 90% — for one component, for specific reasons that do not generalise.
AWS S3 2017: The Blast Radius of a Typo S3 us-east-1 Outage case-study Failure Modes A mistyped command during routine debugging removed far more capacity than intended, and the affected subsystems had not been restarted in years.
AWS: Shuffle Sharding Shuffle Sharding, Virtual Sharding case-study Cell-Based Architecture Assigning each customer a random subset of workers rather than a fixed shard, so that one abusive tenant affects almost nobody else.
AWS: Static Stability Across Availability Zones Static Stability case-study Static Stability AWS designs services to keep working with the capacity they already have when a zone fails, rather than needing the control plane to provision replacements.
Booking.com's Experimentation Platform case-study Business Architecture Booking.com runs over a thousand concurrent experiments and treats the ability to test any change safely as a platform capability rather than a product feature.
Capital One 2019: From SSRF to the Metadata Endpoint IMDSv1 Breach case-study Identity & Access Management A server-side request forgery reached the cloud instance metadata service, obtained temporary credentials, and used an over-permissioned role to read a very large volume of data.
Capital One's Data Centre Exit case-study Enterprise Architecture A major US bank closed all eight of its data centres and moved fully to public cloud, treating governance automation as the enabling technology rather than a constraint.
Cloudflare 2019: One Regex, Global Outage Cloudflare WAF Outage case-study Load Shedding A firewall rule containing a regular expression with catastrophic backtracking consumed CPU across Cloudflare's entire global network within seconds of deployment.
Cloudflare Workers: Isolates Instead of Containers V8 Isolates case-study Edge Functions Running thousands of tenants per process using V8 isolates rather than containers, trading runtime flexibility for near-zero cold start and extreme density.
Cloudflare: Every Server Runs Every Service Homogeneous Edge case-study Cluster Architecture Rather than dedicating machines to roles, Cloudflare runs the full software stack on every server in every location, which turns capacity into a single fungible pool.
Discord's Message Store Migrations case-study Data Architecture Discord moved from MongoDB to Cassandra to ScyllaDB as message volume grew from millions to trillions, each time for a specific and different reason.
Discord: Hot Partitions at Trillions of Messages Discord Message Storage case-study Partitioning & Sharding Discord partitions messages by channel and time bucket, because a single very busy channel would otherwise concentrate load on one partition.
DoorDash Dispatch: Matching as an Optimisation Problem Deep Red, Assignment Optimisation case-study Streaming & Real-Time Data Assigning deliveries greedily to the nearest driver is locally sensible and globally poor, so the assignment is batched and solved as an optimisation.
DoorDash: Decomposing a Python Monolith DoorDash Microservices Migration case-study Service Boundaries DoorDash moved off a Python monolith as growth made deployment risk and scaling limits unmanageable, and used a facade to migrate incrementally.
Dropbox's Move Off S3 Magic Pocket case-study Cost & FinOps Dropbox moved the majority of its file storage off Amazon S3 onto custom infrastructure, reporting savings that its S-1 filing put at roughly $75 million over two years.
eBay's Architectural Generations case-study Legacy Modernization eBay rewrote its core platform several times across its first decade, each time because the previous generation had hit a limit that could not be tuned away.
Equifax 2017: A Known Patch and an Expired Certificate Equifax Breach case-study Supply Chain Security An unpatched framework vulnerability provided entry, and an expired certificate on a monitoring device meant the exfiltration went undetected for months.
Etsy's Continuous Deployment case-study Software Architecture Etsy moved from infrequent, risky releases to dozens of deploys a day, demonstrating that deployment frequency and stability improve together rather than trading off.
Etsy: Deploying Fifty Times a Day Code as Craft case-study Continuous Integration Discipline Etsy demonstrated that very frequent small deployments reduce risk rather than increasing it, and paired it with blameless postmortems.
Figma's Postgres Sharding case-study Data Architecture Figma delayed sharding for years using replicas and vertical partitioning, then sharded Postgres horizontally without downtime using logical shards and a proxy layer.
GitHub 2018: 43 Seconds of Partition, 24 Hours of Recovery GitHub October 2018 Incident case-study Leader Election A 43-second network partition triggered an automated cross-region database failover, and reconciling the resulting divergence took over 24 hours.
Global Map Serving: Precomputed Tiles at Planet Scale Map Tiles, Tile Pyramid case-study Content Delivery Networks Rendering a map view on demand is impossible at global scale, so the world is precomputed into a pyramid of cacheable tiles addressed by coordinate and zoom.
Google Borg to Kubernetes: Learning From an Internal System Borg, Omega case-study Containers Kubernetes was designed with a decade of Borg experience behind it, and its authors have been explicit about which Borg decisions they deliberately did not repeat.
Google Maps and Planetary-Scale Spatial Serving case-study Performance & Capacity Map serving is fast because almost nothing is computed on request — the world is precomputed into a pyramid of tiles, and space is indexed onto a one-dimensional curve.
Google Spanner: Buying Consistency With Time TrueTime, External Consistency case-study Distributed Transactions Spanner achieves globally consistent transactions by bounding clock uncertainty with dedicated hardware and deliberately waiting out that uncertainty on commit.
Google SRE: Error Budgets as a Negotiation Device SRE Error Budget case-study Error Budgets Google resolved the standing conflict between shipping speed and reliability by giving both sides a shared number and a pre-agreed consequence.
Instagram's Early Scaling case-study Architecture Fundamentals Instagram reached tens of millions of users on Django and PostgreSQL with a handful of engineers, by deliberately choosing boring technology and doing the simple thing first.
Knight Capital: $440 Million in 45 Minutes Knight Capital Incident case-study Feature Flags A deployment that reached seven of eight servers, combined with a reused feature flag, activated dormant test code and destroyed the company in three quarters of an hour.
LinkedIn and the Origin of Kafka case-study API & Integration Kafka was built to replace point-to-point data integration between many systems with a single durable log that any system could publish to and any number could read.
LinkedIn Databus: Change Capture as a Product Databus case-study Change Data Capture LinkedIn built a change capture system so that derived stores — search, graph, caches — could stay current without every application dual-writing to them.
LinkedIn Project Inversion: Stopping to Fix the Road Inversion case-study Technical Debt LinkedIn halted feature development for roughly two months to rebuild its deployment and development infrastructure, because the tooling had become the constraint on everything.
LinkedIn: Kafka and the Unified Log The Log, Kafka Origin case-study Event Streaming LinkedIn replaced a tangle of point-to-point data pipelines with a single durable log, turning an O(n²) integration problem into an O(n) one.
Maersk and NotPetya: Recovery from a Single Surviving Copy NotPetya 2017 case-study Backup Strategies A destructive malware outbreak encrypted Maersk's estate globally, and the domain controllers were recovered only because one office had been offline during the attack.
Meta 2021: The Outage That Locked Out Its Own Engineers Facebook October 2021 Outage case-study Routing & BGP A backbone configuration command withdrew Facebook's BGP routes globally, and the tools needed to fix it depended on the network that had just disappeared.
Monzo's Microservice Estate case-study Software Architecture Monzo runs a bank on well over a thousand microservices, and the interesting engineering is in the platform and network isolation that makes that number survivable.
Netflix Open Connect case-study Networking Netflix built its own CDN and placed appliances inside ISP networks, turning the most expensive part of its cost structure into hardware it controls.
Netflix Spinnaker: Automated Canary Analysis Spinnaker, Kayenta case-study Progressive Delivery Netflix made canary analysis a statistical comparison run by a machine rather than an engineer watching a dashboard.
Netflix's Recommendation Architecture case-study Data Architecture Netflix splits personalisation into offline, nearline and online layers so that expensive computation happens ahead of time and the request path stays fast.
Netflix: Chaos Monkey and the Simian Army Simian Army, Chaos Kong case-study Chaos Engineering Netflix deliberately terminated production instances during business hours to force engineers to build for failure rather than hope against it.
Netflix: From Hystrix to Adaptive Concurrency Limits Hystrix, Concurrency Limits case-study Circuit Breakers Netflix's widely-copied circuit breaker library was retired in favour of limits that derive themselves from observed latency, because static thresholds go stale.
Netflix: Regional Evacuation Chaos Kong, Region Failover case-study Multi-Region Architecture Netflix rehearses shifting all traffic out of an entire AWS region, which is what makes the capability real rather than documented.
Pinterest's MySQL Sharding case-study Data Architecture Pinterest sharded MySQL by embedding the shard ID inside every primary key, making any object's location computable from its ID alone with no lookup service.
Prime Video's Move Back to a Monolith Prime Video VQA Rearchitecture case-study Architecture Patterns Amazon Prime Video consolidated a serverless, distributed audio/video monitoring service into a single process and reported a 90% cost reduction — the most-cited example of microservices being the wrong tool.
Retail Peak: Designing for Black Friday Black Friday Readiness case-study Peak Event Readiness A retailer's annual peak can be an order of magnitude above normal, arrives in minutes, and cannot be rescheduled — which makes it a distinct engineering discipline.
Roblox 2021: A Coordination Layer as a Single Point of Failure Roblox 73-Hour Outage case-study Service Discovery A performance problem in the shared service-discovery cluster took the entire platform down for 73 hours, and its novelty made it extremely difficult to diagnose.
Salesforce's Metadata-Driven Multi-Tenancy case-study Data Architecture Salesforce serves every customer from shared infrastructure with a single physical schema, storing customer-specific data structures as metadata rather than as separate tables.
Shopify's Pods and Modular Monolith case-study Architecture Patterns Shopify handles Black Friday scale with isolated pods — complete stacks each serving a subset of merchants — while keeping the application itself a deliberately modular monolith.
Shopify: A Monolith That Scaled Packwerk, Shopify Pods case-study Modular Monolith Shopify kept its Rails monolith and invested in enforced internal boundaries and horizontal sharding, rather than decomposing into microservices.
Slack 2021: When Autoscaling Cannot Keep Up Slack January 2021 Outage case-study Autoscaling The first Monday back after the holidays produced a traffic ramp that outpaced the scaling behaviour of a managed network component, and the degradation cascaded.
Slack Flannel: Caching at the Edge for a Chat Client Flannel case-study Caching Strategies Slack pushed user and channel metadata into an application-aware edge cache because clients were downloading enormous amounts of it on every connection.
Slack's Cellular Migration case-study Cloud Architecture After repeated availability-zone-level incidents, Slack rebuilt its infrastructure into per-zone cells with the ability to drain traffic away from a failing zone in minutes.
Southwest 2022: Legacy Software Meeting Its Design Limits Southwest Airlines Meltdown case-study Modernisation Business Case A crew scheduling system that worked adequately for routine disruption could not cope with a large-scale one, and the airline cancelled roughly 16,700 flights.
Spotify Backstage: The Catalogue as a Product Backstage case-study Internal Developer Platform Spotify built a developer portal that unified service discovery, documentation, templates and tooling behind one interface, and open-sourced it.
Spotify Discover Weekly: Three Models, One Playlist Discover Weekly case-study ML Platform Spotify combined collaborative filtering, natural language processing and raw audio analysis because each covers the others' blind spots.