Terminology
2185 terms, tools, patterns and metrics an architect is expected to use precisely. Each one gets a short explanation of what it is, and — where it matters — what it is commonly confused with. Search filters as you type; the column headers sort.
All areas2185
Architecture Fundamentals77
Distributed Systems107
Data Architecture110
Cloud Architecture89
Networking88
API & Integration Architecture82
Reliability & Resilience75
Observability70
Performance & Capacity Engineering72
Security Architecture81
Cost Architecture & FinOps66
Business Architecture67
Architecture Communication67
Enterprise Architecture66
Legacy Modernization66
AI-Era Architecture69
Software Architecture & Engineering71
Architecture Patterns71
Architecture Decision-Making63
The Architect's Meta-Skills61
Delivery & Release Engineering66
Platform Engineering & Developer Experience70
Testing & Quality Architecture66
Data Platform Architecture64
Streaming & Real-Time Data69
Data Governance & Semantics69
Frontend & Experience Architecture66
Edge, Mobile & IoT68
Regulatory & Data Protection Architecture65
Assurance, Audit & Model Risk64
77 terms shown.
| Term | Kind | Topic | What it is |
|---|---|---|---|
| Airbnb Minerva: One Definition of a Metric Minerva Metrics Platform | case-study | Semantic Layer | Airbnb built a central metrics platform because the same business question was producing different answers depending on which dashboard you opened. |
| Airbnb's Service-Oriented Migration | case-study | Legacy Modernization | Airbnb decomposed a large Rails monolith by first extracting a unified data-access layer, so that services were built on owned data rather than on shared database tables. |
| Airbnb: Monorail to Service-Oriented Architecture Airbnb SOA Migration | case-study | Application Decomposition | Airbnb decomposed a large Rails monolith into services, and the difficult part was data ownership rather than code extraction. |
| Airflow: A Scheduler Born From Pipeline Sprawl Apache Airflow | case-study | Workflow Schedulers | Airbnb built Airflow because cron cannot express dependencies, and a data platform's failures are mostly dependency failures. |
| Amazon 2002: The API Mandate Bezos Mandate | case-study | API & Integration | A directive that all teams expose functionality only through service interfaces, with no back doors, is widely credited with making both Amazon's architecture and AWS possible. |
| Amazon Dynamo: Always Writeable Dynamo Paper | case-study | Consistency Models | Amazon chose an always-writeable shopping cart with application-level conflict resolution, accepting merge complexity to guarantee that "add to cart" never fails. |
| Amazon Prime Video: Serverless Back to a Monolith Prime Video Audio/Video Monitoring | case-study | Monolith vs Microservices | A distributed serverless pipeline was consolidated into a single process, reducing cost by over 90% — for one component, for specific reasons that do not generalise. |
| AWS S3 2017: The Blast Radius of a Typo S3 us-east-1 Outage | case-study | Failure Modes | A mistyped command during routine debugging removed far more capacity than intended, and the affected subsystems had not been restarted in years. |
| AWS: Shuffle Sharding Shuffle Sharding, Virtual Sharding | case-study | Cell-Based Architecture | Assigning each customer a random subset of workers rather than a fixed shard, so that one abusive tenant affects almost nobody else. |
| AWS: Static Stability Across Availability Zones Static Stability | case-study | Static Stability | AWS designs services to keep working with the capacity they already have when a zone fails, rather than needing the control plane to provision replacements. |
| Booking.com's Experimentation Platform | case-study | Business Architecture | Booking.com runs over a thousand concurrent experiments and treats the ability to test any change safely as a platform capability rather than a product feature. |
| Capital One 2019: From SSRF to the Metadata Endpoint IMDSv1 Breach | case-study | Identity & Access Management | A server-side request forgery reached the cloud instance metadata service, obtained temporary credentials, and used an over-permissioned role to read a very large volume of data. |
| Capital One's Data Centre Exit | case-study | Enterprise Architecture | A major US bank closed all eight of its data centres and moved fully to public cloud, treating governance automation as the enabling technology rather than a constraint. |
| Cloudflare 2019: One Regex, Global Outage Cloudflare WAF Outage | case-study | Load Shedding | A firewall rule containing a regular expression with catastrophic backtracking consumed CPU across Cloudflare's entire global network within seconds of deployment. |
| Cloudflare Workers: Isolates Instead of Containers V8 Isolates | case-study | Edge Functions | Running thousands of tenants per process using V8 isolates rather than containers, trading runtime flexibility for near-zero cold start and extreme density. |
| Cloudflare: Every Server Runs Every Service Homogeneous Edge | case-study | Cluster Architecture | Rather than dedicating machines to roles, Cloudflare runs the full software stack on every server in every location, which turns capacity into a single fungible pool. |
| Discord's Message Store Migrations | case-study | Data Architecture | Discord moved from MongoDB to Cassandra to ScyllaDB as message volume grew from millions to trillions, each time for a specific and different reason. |
| Discord: Hot Partitions at Trillions of Messages Discord Message Storage | case-study | Partitioning & Sharding | Discord partitions messages by channel and time bucket, because a single very busy channel would otherwise concentrate load on one partition. |
| DoorDash Dispatch: Matching as an Optimisation Problem Deep Red, Assignment Optimisation | case-study | Streaming & Real-Time Data | Assigning deliveries greedily to the nearest driver is locally sensible and globally poor, so the assignment is batched and solved as an optimisation. |
| DoorDash: Decomposing a Python Monolith DoorDash Microservices Migration | case-study | Service Boundaries | DoorDash moved off a Python monolith as growth made deployment risk and scaling limits unmanageable, and used a facade to migrate incrementally. |
| Dropbox's Move Off S3 Magic Pocket | case-study | Cost & FinOps | Dropbox moved the majority of its file storage off Amazon S3 onto custom infrastructure, reporting savings that its S-1 filing put at roughly $75 million over two years. |
| eBay's Architectural Generations | case-study | Legacy Modernization | eBay rewrote its core platform several times across its first decade, each time because the previous generation had hit a limit that could not be tuned away. |
| Equifax 2017: A Known Patch and an Expired Certificate Equifax Breach | case-study | Supply Chain Security | An unpatched framework vulnerability provided entry, and an expired certificate on a monitoring device meant the exfiltration went undetected for months. |
| Etsy's Continuous Deployment | case-study | Software Architecture | Etsy moved from infrequent, risky releases to dozens of deploys a day, demonstrating that deployment frequency and stability improve together rather than trading off. |
| Etsy: Deploying Fifty Times a Day Code as Craft | case-study | Continuous Integration Discipline | Etsy demonstrated that very frequent small deployments reduce risk rather than increasing it, and paired it with blameless postmortems. |
| Figma's Postgres Sharding | case-study | Data Architecture | Figma delayed sharding for years using replicas and vertical partitioning, then sharded Postgres horizontally without downtime using logical shards and a proxy layer. |
| GitHub 2018: 43 Seconds of Partition, 24 Hours of Recovery GitHub October 2018 Incident | case-study | Leader Election | A 43-second network partition triggered an automated cross-region database failover, and reconciling the resulting divergence took over 24 hours. |
| Global Map Serving: Precomputed Tiles at Planet Scale Map Tiles, Tile Pyramid | case-study | Content Delivery Networks | Rendering a map view on demand is impossible at global scale, so the world is precomputed into a pyramid of cacheable tiles addressed by coordinate and zoom. |
| Google Borg to Kubernetes: Learning From an Internal System Borg, Omega | case-study | Containers | Kubernetes was designed with a decade of Borg experience behind it, and its authors have been explicit about which Borg decisions they deliberately did not repeat. |
| Google Maps and Planetary-Scale Spatial Serving | case-study | Performance & Capacity | Map serving is fast because almost nothing is computed on request — the world is precomputed into a pyramid of tiles, and space is indexed onto a one-dimensional curve. |
| Google Spanner: Buying Consistency With Time TrueTime, External Consistency | case-study | Distributed Transactions | Spanner achieves globally consistent transactions by bounding clock uncertainty with dedicated hardware and deliberately waiting out that uncertainty on commit. |
| Google SRE: Error Budgets as a Negotiation Device SRE Error Budget | case-study | Error Budgets | Google resolved the standing conflict between shipping speed and reliability by giving both sides a shared number and a pre-agreed consequence. |
| Instagram's Early Scaling | case-study | Architecture Fundamentals | Instagram reached tens of millions of users on Django and PostgreSQL with a handful of engineers, by deliberately choosing boring technology and doing the simple thing first. |
| Knight Capital: $440 Million in 45 Minutes Knight Capital Incident | case-study | Feature Flags | A deployment that reached seven of eight servers, combined with a reused feature flag, activated dormant test code and destroyed the company in three quarters of an hour. |
| LinkedIn and the Origin of Kafka | case-study | API & Integration | Kafka was built to replace point-to-point data integration between many systems with a single durable log that any system could publish to and any number could read. |
| LinkedIn Databus: Change Capture as a Product Databus | case-study | Change Data Capture | LinkedIn built a change capture system so that derived stores — search, graph, caches — could stay current without every application dual-writing to them. |
| LinkedIn Project Inversion: Stopping to Fix the Road Inversion | case-study | Technical Debt | LinkedIn halted feature development for roughly two months to rebuild its deployment and development infrastructure, because the tooling had become the constraint on everything. |
| LinkedIn: Kafka and the Unified Log The Log, Kafka Origin | case-study | Event Streaming | LinkedIn replaced a tangle of point-to-point data pipelines with a single durable log, turning an O(n²) integration problem into an O(n) one. |
| Maersk and NotPetya: Recovery from a Single Surviving Copy NotPetya 2017 | case-study | Backup Strategies | A destructive malware outbreak encrypted Maersk's estate globally, and the domain controllers were recovered only because one office had been offline during the attack. |
| Meta 2021: The Outage That Locked Out Its Own Engineers Facebook October 2021 Outage | case-study | Routing & BGP | A backbone configuration command withdrew Facebook's BGP routes globally, and the tools needed to fix it depended on the network that had just disappeared. |
| Monzo's Microservice Estate | case-study | Software Architecture | Monzo runs a bank on well over a thousand microservices, and the interesting engineering is in the platform and network isolation that makes that number survivable. |
| Netflix Open Connect | case-study | Networking | Netflix built its own CDN and placed appliances inside ISP networks, turning the most expensive part of its cost structure into hardware it controls. |
| Netflix Spinnaker: Automated Canary Analysis Spinnaker, Kayenta | case-study | Progressive Delivery | Netflix made canary analysis a statistical comparison run by a machine rather than an engineer watching a dashboard. |
| Netflix's Recommendation Architecture | case-study | Data Architecture | Netflix splits personalisation into offline, nearline and online layers so that expensive computation happens ahead of time and the request path stays fast. |
| Netflix: Chaos Monkey and the Simian Army Simian Army, Chaos Kong | case-study | Chaos Engineering | Netflix deliberately terminated production instances during business hours to force engineers to build for failure rather than hope against it. |
| Netflix: From Hystrix to Adaptive Concurrency Limits Hystrix, Concurrency Limits | case-study | Circuit Breakers | Netflix's widely-copied circuit breaker library was retired in favour of limits that derive themselves from observed latency, because static thresholds go stale. |
| Netflix: Regional Evacuation Chaos Kong, Region Failover | case-study | Multi-Region Architecture | Netflix rehearses shifting all traffic out of an entire AWS region, which is what makes the capability real rather than documented. |
| Pinterest's MySQL Sharding | case-study | Data Architecture | Pinterest sharded MySQL by embedding the shard ID inside every primary key, making any object's location computable from its ID alone with no lookup service. |
| Prime Video's Move Back to a Monolith Prime Video VQA Rearchitecture | case-study | Architecture Patterns | Amazon Prime Video consolidated a serverless, distributed audio/video monitoring service into a single process and reported a 90% cost reduction — the most-cited example of microservices being the wrong tool. |
| Retail Peak: Designing for Black Friday Black Friday Readiness | case-study | Peak Event Readiness | A retailer's annual peak can be an order of magnitude above normal, arrives in minutes, and cannot be rescheduled — which makes it a distinct engineering discipline. |
| Roblox 2021: A Coordination Layer as a Single Point of Failure Roblox 73-Hour Outage | case-study | Service Discovery | A performance problem in the shared service-discovery cluster took the entire platform down for 73 hours, and its novelty made it extremely difficult to diagnose. |
| Salesforce's Metadata-Driven Multi-Tenancy | case-study | Data Architecture | Salesforce serves every customer from shared infrastructure with a single physical schema, storing customer-specific data structures as metadata rather than as separate tables. |
| Shopify's Pods and Modular Monolith | case-study | Architecture Patterns | Shopify handles Black Friday scale with isolated pods — complete stacks each serving a subset of merchants — while keeping the application itself a deliberately modular monolith. |
| Shopify: A Monolith That Scaled Packwerk, Shopify Pods | case-study | Modular Monolith | Shopify kept its Rails monolith and invested in enforced internal boundaries and horizontal sharding, rather than decomposing into microservices. |
| Slack 2021: When Autoscaling Cannot Keep Up Slack January 2021 Outage | case-study | Autoscaling | The first Monday back after the holidays produced a traffic ramp that outpaced the scaling behaviour of a managed network component, and the degradation cascaded. |
| Slack Flannel: Caching at the Edge for a Chat Client Flannel | case-study | Caching Strategies | Slack pushed user and channel metadata into an application-aware edge cache because clients were downloading enormous amounts of it on every connection. |
| Slack's Cellular Migration | case-study | Cloud Architecture | After repeated availability-zone-level incidents, Slack rebuilt its infrastructure into per-zone cells with the ability to drain traffic away from a failing zone in minutes. |
| Southwest 2022: Legacy Software Meeting Its Design Limits Southwest Airlines Meltdown | case-study | Modernisation Business Case | A crew scheduling system that worked adequately for routine disruption could not cope with a large-scale one, and the airline cancelled roughly 16,700 flights. |
| Spotify Backstage: The Catalogue as a Product Backstage | case-study | Internal Developer Platform | Spotify built a developer portal that unified service discovery, documentation, templates and tooling behind one interface, and open-sourced it. |
| Spotify Discover Weekly: Three Models, One Playlist Discover Weekly | case-study | ML Platform | Spotify combined collaborative filtering, natural language processing and raw audio analysis because each covers the others' blind spots. |
Nothing on this page matches. Search the whole glossary.