Terminology
2185 terms, tools, patterns and metrics an architect is expected to use precisely. Each one gets a short explanation of what it is, and — where it matters — what it is commonly confused with. Search filters as you type; the column headers sort.
All areas2185
Architecture Fundamentals77
Distributed Systems107
Data Architecture110
Cloud Architecture89
Networking88
API & Integration Architecture82
Reliability & Resilience75
Observability70
Performance & Capacity Engineering72
Security Architecture81
Cost Architecture & FinOps66
Business Architecture67
Architecture Communication67
Enterprise Architecture66
Legacy Modernization66
AI-Era Architecture69
Software Architecture & Engineering71
Architecture Patterns71
Architecture Decision-Making63
The Architect's Meta-Skills61
Delivery & Release Engineering66
Platform Engineering & Developer Experience70
Testing & Quality Architecture66
Data Platform Architecture64
Streaming & Real-Time Data69
Data Governance & Semantics69
Frontend & Experience Architecture66
Edge, Mobile & IoT68
Regulatory & Data Protection Architecture65
Assurance, Audit & Model Risk64
75 terms shown.
| Term | Kind | Topic | What it is |
|---|---|---|---|
| Aspirational and Achievable SLO | practice | SLI, SLO & SLA | The distinction between the reliability a team wishes for and the reliability its current architecture and dependencies can actually deliver. |
| Availability Arithmetic | concept | Availability Mathematics | Multiplying dependency availabilities in series and combining redundant components in parallel to derive a system's achievable availability. |
| Availability Calculation | concept | Reliability & Resilience | Deriving a system's availability from its components, remembering that dependencies in series multiply. |
| Availability Composition | concept | Availability Mathematics | How the availability of a system follows from its dependencies — multiplying for serial dependencies and improving sharply for redundant ones. |
| AWS: Static Stability Across Availability Zones Static Stability | case-study | Static Stability | AWS designs services to keep working with the capacity they already have when a zone fails, rather than needing the control plane to provision replacements. |
| Blameless Postmortem Learning Review, Incident Retrospective, Non-Attributive Analysis | practice | Postmortems | An incident analysis structured so that participants can disclose what actually happened without personal risk - which is a technical requirement for accurate information, not a cultural courtesy. |
| Blameless Postmortem | practice | Reliability & Resilience | An incident review that seeks the systemic conditions that made a failure possible, explicitly excluding individual fault. |
| Blast Radius Failure Scope, Impact Radius, Containment Boundary | concept | Fault Isolation | The set of users, tenants, regions or services that a single failure can affect - the quantity that distinguishes an incident from a catastrophe, and the one that architecture can most directly control. |
| Blast Radius Analysis | practice | Fault Isolation | Determining, for each component, exactly what fails and who is affected when it fails completely. |
| Blast Radius Staging Graduated Rollout, Progressive Shedding | practice | Failover | Applying a change or a traffic shift in increasing increments gated on health, so that a wrong decision affects a bounded population before it affects everyone. |
| Capacity Headroom | concept | Capacity Planning | The deliberate gap between current load and maximum capacity, sized to absorb growth, spikes and the loss of a failure domain. |
| Capacity Planning | practice | Reliability & Resilience | Deciding in advance how much capacity will be needed, given growth, seasonality and failure scenarios, and ensuring it can be there in time. |
| Chaos Engineering | practice | Reliability & Resilience | Deliberately injecting failure into a system to discover, before an incident does, which of your resilience assumptions are false. |
| Chaos Engineering in Practice | practice | Chaos Engineering | Deliberately injecting failure into production to verify that resilience mechanisms work — an experiment with a hypothesis, not vandalism. |
| Concurrency Limiting | pattern | Fault Isolation | Bounding the number of simultaneous in-flight operations so that overload produces fast rejection rather than resource exhaustion. |
| Contributing Factors Analysis | practice | Postmortems | Identifying the multiple conditions that combined to produce an incident, in place of searching for a single root cause. |
| Correlated Degradation Demand-Coupled Failure, Peak-Correlated Dependency Risk | concept | Degradation Modes | The condition where a dependency's reliability worsens for the same reason that user demand rises, so the system is least capable exactly when it is most needed. |
| Correlated Failure Common-Mode Failure, Shared Fate | concept | Redundancy | Failures that are not independent, so redundancy multiplies far less than the arithmetic promises - usually because replicas share code, configuration or a deployment. |
| Dark Failover Shadow Failover, Non-Serving Failover Rehearsal | practice | DR Testing | Bringing a standby environment to full serving readiness and exercising it with synthetic or internal traffic while production continues untouched, so recovery is measured without an outage being risked. |
| Degradation Ladder Kill Switch List, Feature Shedding Order, Brownout Plan | practice | Degradation Modes | An ordered, pre-agreed list of features that can be disabled under stress, each with a switch that works without deployment, a named owner, a trigger condition and a rehearsed fallback. |
| Degradation Mode | concept | Degradation Modes | A defined, intentional reduced state of service that the system enters under specified conditions, with known behaviour and known exit criteria. |
| Degradation Modes | concept | Degradation Modes | Explicitly designed operating modes below "fully working" — declared, testable, and switchable rather than emergent. |
| Disaster Recovery DR | practice | Reliability & Resilience | The plan and capability for restoring service after an event that takes out a whole site, region or system. |
| Durable Execution Workflow Persistence, Resumable Orchestration | pattern | Reliability & Resilience | Persisting a workflow's progress outside the process executing it, so that worker restarts, deployments and crashes resume from the last completed step rather than losing position. |
| Error Budget Reliability Budget, SLO Budget, Failure Allowance | practice | Error Budgets | The permitted amount of unreliability implied by an SLO, used as an automatic arbiter between shipping features and improving reliability - and, when unspent, as evidence that the system is more reliable than … |
| Error Budget | metric | Reliability & Resilience | The amount of unreliability an SLO permits, treated as a resource that feature velocity spends. |
| Error Budget Policy | practice | Error Budgets | The pre-agreed consequences of exhausting an error budget, which is what turns an SLO from a metric into a decision-making mechanism. |
| Failover | pattern | Failover | Switching to a standby when the primary fails — where detection, fencing and the decision to automate are harder than the switch itself. |
| Failover Orchestration | pattern | Failover | The sequence of detection, decision, promotion and traffic redirection that moves service from a failed component to a healthy one. |
| Failover Propagation Delay Client Convergence Time, Failover Tail | metric | Failover | The interval between a failover completing on the server side and the last client actually using the new location, which is usually the larger part of the recovery time users experience. |
| Failure Injection Testing | practice | Resilience Testing | Deliberately introducing faults into a system under test to verify that timeouts, retries, fallbacks and circuit breakers behave as designed. |
| Fallback Independence | practice | Graceful Degradation | The requirement that a degraded path fail independently of the primary - and be exercised continuously, because fallbacks that are never run do not work. |
| Fault Domain | concept | Fault Isolation | A boundary within which a single failure is contained, defined by the infrastructure and dependencies that components inside it share. |
| Feature Criticality Tiering | practice | Graceful Degradation | Classifying product functionality by whether it must work, should work, or can be dropped, so degradation decisions are made in advance by the business. |
| Game Day | practice | Reliability & Resilience | A scheduled exercise in which a failure is deliberately introduced and the team responds as though it were real, to test the system and the response together. |
| Game Days | practice | Game Days | Scheduled exercises where a team responds to a simulated or injected failure, testing the humans and the process as much as the system. |
| Google SRE: Error Budgets as a Negotiation Device SRE Error Budget | case-study | Error Budgets | Google resolved the standing conflict between shipping speed and reliability by giving both sides a shared number and a pre-agreed consequence. |
| Graceful Degradation in Practice | pattern | Graceful Degradation | Deliberately reducing functionality to preserve the core when a dependency fails — a product decision expressed in architecture. |
| Health Check Depth Shallow vs Deep Health Check, Dependency Health Check | concept | Failover | How much of a service's dependency graph a health check consults, which determines whether the check reports an independent fault or a shared one that every instance will report at the same moment. |
| Incident Command Incident Command System, ICS | practice | Reliability & Resilience | Assigning explicit roles during an incident — commander, operations lead, communications lead, scribe — so coordination does not compete with diagnosis. |
| Incident Severity Levels | practice | Incident Management | A small, agreed scale of incident severity that determines response, escalation and communication without requiring debate during the event. |
| Journey-Level SLO Per-Journey Availability, User-Facing SLI | practice | SLI, SLO & SLA | Setting reliability targets per user journey rather than per platform or per service, because consequence differs by journey and only a journey-level measurement reflects what a user experienced. |
| Load-Coupled Fault Injection Fault Injection Under Load, Contention Testing | practice | Chaos Engineering | Injecting dependency faults while the system is at realistic peak load, because timeouts, pool sizes and breaker thresholds only reveal whether they are correct under contention. |
| N+1 and 2N Redundancy | pattern | Redundancy | Provisioning one spare beyond required capacity versus provisioning double, and the failure assumptions each encodes. |
| Netflix: Chaos Monkey and the Simian Army Simian Army, Chaos Kong | case-study | Chaos Engineering | Netflix deliberately terminated production instances during business hours to force engineers to build for failure rather than hope against it. |
| Nines Table | concept | Availability Mathematics | The mapping between availability percentages and permitted downtime, and the cost curve that comes with it. |
| On-Call | practice | On-Call | The rotation that responds to production problems — a system whose health is measured by whether the people in it can sustain being in it. |
| Page Budget Pages Per Shift, Alert Budget | metric | On-Call | An explicit ceiling on how many pages a shift may generate, treated as a limit the team manages against rather than as an outcome it observes. |
| Postmortems Incident Review, Blameless Postmortem | practice | Postmortems | Structured learning from an incident, conducted so that the truth is obtainable — which requires that telling it is safe. |
| Pre-Provisioned Capacity | pattern | Static Stability | Running capacity that is already in place to absorb a failure, rather than depending on a control plane to create it during the failure. |
| Recovery Cost Per Unit Remediation Cost, Recoverability, Time-to-Repair Per Node | concept | Fault Isolation | The effort required to restore one affected instance, device or tenant - the multiplier that decides whether a wide blast radius is an incident or a catastrophe, and the factor most often left out of risk asse… |
| Recovery Point Objective RPO | metric | RTO & RPO | The maximum acceptable data loss measured as a duration, determining the replication and backup strategy. |
| Recovery Time Objective RTO | metric | RTO & RPO | The maximum acceptable duration between a failure and restored service, agreed with the business rather than chosen by engineering. |
| Redundancy | concept | Redundancy | Multiple instances of a component so that one failing does not fail the system — valuable exactly to the extent the failures are uncorrelated. |
| Reliability Headroom | metric | Capacity Planning | The gap between provisioned capacity and the load that would be carried after the largest planned failure, measured under peak conditions. |
| Request-Based and Window-Based SLI | metric | SLI, SLO & SLA | Two ways of computing the same reliability target — counting good events, or counting good time windows — which produce materially different numbers. |
| Resilience Testing | practice | Resilience Testing | Verifying that failure-handling behaves as designed — a category of testing distinct from functional and load testing, and usually absent. |
| Restore Granularity Recovery Unit, Restore Scope | concept | RTO & RPO | The smallest unit a recovery procedure can return without disturbing everything stored alongside it, which sets the real recovery time for any incident that affects a subset of the data. |
| Restore Verification | practice | DR Testing | Periodically performing a real restore from backup and validating the result, as the only evidence that a recovery capability exists. |
| RTO and RPO | metric | RTO & RPO | How long recovery may take, and how much data may be lost — the two numbers from which every disaster recovery design follows. |
Nothing on this page matches. Search the whole glossary.