Term Kind Topic What it is
Service Registry tool Service Discovery The database of currently available service instances and their addresses, maintained by registration and pruned by health checking.
Service Scaffolding tool Service Templates Generating a new service with observability, CI, security and ownership already wired in, so the standard is the default rather than a checklist.
Service Stub Fidelity concept Service Virtualisation How faithfully a stand-in dependency reproduces the real one's behaviour — including its errors, latency and limits — which bounds what testing against it proves.
Service Template Scaffold, Cookiecutter Template, Software Template tool Service Templates A generator that creates a new service already wired into the organisation's pipeline, telemetry, security and operational conventions.
Service Virtualisation API Simulation, Mock Server tool Service Virtualisation A stand-in for a real dependency that responds with realistic data, latency and error behaviour, so a service can be tested without it.
Serving Latency Budget concept Real-Time Serving Layer The end-to-end time from an event occurring to its effect being queryable, allocated across ingest, processing and serving.
Serving Path Divergence Per-Pool Quality Drift, Heterogeneous Fleet Skew concept AI Observability One configuration in a fleet of otherwise identical inference paths behaving differently from its peers - detectable by comparing paths against each other, and invisible to any metric averaged across them.
Severity Levels SEV Levels practice Reliability & Resilience A predefined scale of incident impact that determines who is woken, how fast, and what process applies.
Shadow Comparison pattern Parallel Run Running a new implementation alongside the old on real traffic, comparing outputs without the new system affecting users.
Shadow Planning Plan Shadowing, Shadow Query Planning, Dry-Run Routing practice Refactoring Running a new routing or planning layer over live production traffic without serving its results, so the work it cannot handle is enumerated from real usage rather than from reading code.
Shard Key Partition Key practice Partitioning & Sharding The attribute deciding which partition a row belongs to - the single most consequential and least reversible choice in a partitioned data architecture.
Shard Key Selection practice Sharding Patterns Choosing the attribute that determines a record's partition, which fixes the system's distribution, query patterns and future flexibility.
Sharding Horizontal Partitioning pattern Data Architecture Splitting one dataset across multiple independent databases by a partition key, so that each holds a disjoint subset.
Sharding in Practice Horizontal Partitioning pattern Partitioning & Sharding Splitting data across independent stores, how to choose the key, and why resharding is the operation nobody plans for.
Sharding Patterns Horizontal Partitioning pattern Sharding Patterns Splitting data across independent stores by a key, and the choice of that key — which is the decision that will define the system for years.
Shared Database Integration Integration Database concept Legacy Integration Two or more applications reading and writing the same database directly — the most damaging integration pattern and the hardest to unwind.
Shared Kernel Shared Model, Common Domain Library pattern Bounded Contexts A deliberately shared subset of the domain model between two contexts, which removes translation cost by accepting a shared release cadence - the one context relationship that couples teams on purpose.
Shared Responsibility Model concept Managed Services The division of duties between provider and customer, which shifts with the service model and is routinely misunderstood in the customer's disfavour.
Shared Singleton Skew Duplicate Framework Instance, Singleton Version Mismatch concept Module Federation The failure class created when independently built front-end bundles expect different versions of a library that must exist once per page, producing either incompatible reuse or two isolated copies that share …
Shared-Cause Analysis What Is Shared, Redundancy Audit practice Redundancy Auditing a redundant design by asking what every replica has in common, because redundancy protects against independent failure and not against anything shared by all copies.
Shared-Service Commons Free-at-Point-of-Use Platform, Uncharged Internal Platform concept Showback & Chargeback The failure mode where cost allocation prices each team's own resources but leaves internal platforms free to use, so teams move cost onto the shared estate instead of removing it and total spend rises while e…
Shared-Team Queue Central Team Bottleneck, Review Queue concept Operating Models The queueing behaviour of any central team that must approve or perform work for many others, where waiting time rises sharply as utilisation approaches capacity and teams begin routing around the control.
Shell Personalisation Cache the Shared, Personalise the Edge, Hole Punching pattern Rendering Strategies Generating and caching the page that is identical for everyone, then applying per-user variation at the edge or on the client - preserving the cache hit rate that personalising the whole page destroys.
Shift-Left Security practice Security Testing in the Pipeline Moving security checks earlier so findings arrive while the author still has context, on the condition that the signal-to-noise ratio justifies it.
Shopify's Pods and Modular Monolith case-study Architecture Patterns Shopify handles Black Friday scale with isolated pods — complete stacks each serving a subset of merchants — while keeping the application itself a deliberately modular monolith.
Shopify: A Monolith That Scaled Packwerk, Shopify Pods case-study Modular Monolith Shopify kept its Rails monolith and invested in enforced internal boundaries and horizontal sharding, rather than decomposing into microservices.
Showback practice Platform Funding Reporting each team's share of platform and infrastructure cost without actually charging it, to change behaviour without creating a market.
Showback and Chargeback practice Showback & Chargeback Showing teams their costs, or actually billing them — two different mechanisms with different incentive effects and different failure modes.
Shuffle Sharding Virtual Sharding, Randomised Subset Assignment pattern Bulkheads & Isolation Assigning each tenant a random subset of the available capacity rather than a single shard, so that any two tenants rarely share their entire subset and one tenant's failure affects almost nobody completely.
Side-Effect Isolation Synthetic Traffic Containment, Shadow Write Isolation, Test Traffic Tagging practice Testing in Production Ensuring that test, synthetic or shadowed traffic in production cannot produce real-world effects - the control that separates professional production testing from an incident caused deliberately.
Sidecar pattern Architecture Patterns Deploying a helper process alongside the main application in the same unit, to supply cross-cutting behaviour without changing the application.
Sidecar Interception Boundary Mesh Coverage Boundary, Proxy Capture Scope concept Sidecar & Ambassador The precise set of traffic a sidecar proxy actually sees and can secure, which is narrower than "all traffic from this pod" and is what decides whether a mesh's guarantees apply to a given call.
Sidecar Resource Overhead concept Sidecar & Ambassador The cumulative CPU, memory and latency cost of running helper containers alongside every application container.
Significance Trigger Review Trigger, Architecturally Significant Change practice Architecture Review Boards A pre-defined, published condition that requires architectural review - so that governance covers what matters and does not become a queue in front of everything else.
Significant-Minority Trigger Defined Escalation Criteria, When Humans Review practice Deployment Gates A stated rule defining which changes require human review, so the reservation of manual approval is a rule rather than a judgement made under time pressure.
Signing Boundary Policy Policy at the HSM, Constrained Signing pattern Key Management Enforcing transaction policy at the hardware signing boundary rather than in application code, so that a compromised application cannot obtain an arbitrary signature.
Silent Coercion Type Coercion on Ingest, Quiet Data Corruption concept Ingestion Patterns An ingestion pipeline converting a value it did not expect rather than rejecting it - the worst available failure mode, because the load succeeds and the data is wrong downstream with nothing indicating it.
Silent Control Failure Dormant Control, Control Decay concept Control Design vs Operation A control that has stopped operating while still reporting success, so the organisation keeps making decisions on the assumption that it is protecting them.
Silent Data Failure No Error, Plausibly Wrong, Quiet Breakage concept Data Observability A data problem that produces no error - the pipeline succeeded, the schema was valid, the numbers were plausible - which is why error-based monitoring cannot detect it and why it persists for weeks.
Silent Quality Regression Payload-Level Degradation, Unasserted Output Failure concept Failure Thinking A defect that leaves status codes, latency and throughput normal while making the content of responses worse - so every envelope-level alert stays green and the only available signal is statistical.
Single-Credential Test Can One Credential Do Both, Segregation Reality Check practice Segregation of Duties Asking whether any single credential - including a database administrator, a root account or a deployment pipeline - can both initiate and approve a movement of value, which distinguishes a real segregation co…
Six Rs Decision Migration Disposition practice The Six Rs Classifying each application into rehost, replatform, refactor, repurchase, retire or retain, based on value and effort rather than on preference.
Six Rs of Migration 6 Rs, Migration Strategies concept Legacy Modernization The standard menu of options for each application in a migration — rehost, replatform, refactor, repurchase, retire, retain.
SLA Measurement Point Measurement Boundary, Observation Point concept Negotiation The place in the request path where an availability or latency commitment is observed, which decides what the number means and is where nearly every SLA dispute actually originates.
Slack 2021: When Autoscaling Cannot Keep Up Slack January 2021 Outage case-study Autoscaling The first Monday back after the holidays produced a traffic ramp that outpaced the scaling behaviour of a managed network component, and the degradation cascaded.
Slack Flannel: Caching at the Edge for a Chat Client Flannel case-study Caching Strategies Slack pushed user and channel metadata into an application-aware edge cache because clients were downloading enormous amounts of it on every connection.
Slack's Cellular Migration case-study Cloud Architecture After repeated availability-zone-level incidents, Slack rebuilt its infrastructure into per-zone cells with the ability to drain traffic away from a failing zone in minutes.
SLI Measurement Point SLI Vantage Point, Indicator Measurement Location practice SLO Monitoring The deliberate choice of where in the request path an indicator is computed, which determines which failures the objective can see at all.
SLI, SLO and SLA metric SLI, SLO & SLA The measurement, the internal target, and the external contract — with the error budget as the mechanism that makes the target consequential.
Slip Cost per Week Weekly Delay Cost, Delay Cost per Week metric Time to Market The money a launch loses for each week it is late - and whether that money is deferred or destroyed - which is the only form in which time to market can be traded against design quality.
Slowly Changing Dimension SCD, Type 2 Dimension pattern Slowly Changing Dimensions A modelling technique for handling attributes that change over time, where the choice determines whether historical reporting stays correct.
Small File Problem concept File Formats & Compaction The severe query degradation caused by a table stored as very many tiny files, where per-file overhead dominates actual data reading.
Small Model Routing Model Cascade, Tiered Inference practice Model Selection Sending each request to the smallest model that can handle it, escalating to a larger one only when needed.
Snapshot Expiry Snapshot Retention, Metadata Expiry practice Open Table Formats The scheduled removal of a table's old snapshots and the data files only those snapshots referenced, which is what stops time travel from turning every rewritten row into permanent storage.
Snapshot-Stream Convergence Backfill Plus Change Stream, Initial Load Ordering pattern CDC Pipeline Design Starting the change stream before taking the snapshot, so that changes occurring during the backfill are captured rather than silently lost.
Snapshotting pattern Event Sourcing Periodically storing an aggregate's computed state so it can be loaded without replaying its entire event history.
Soak Period Bake Time, Observation Window practice Multi-Region Rollout The deliberate waiting time between rollout stages, long enough to surface problems that do not appear immediately - and the first thing compressed under delivery pressure.
Soak Test practice Soak Testing Running sustained realistic load for hours or days to expose defects that accumulate over time rather than appearing under peak load.
Soak Testing practice Soak Testing Running at sustained realistic load for hours or days to find the failures that only appear with time.
Software Bill of Materials SBOM tool Supply-Chain Provenance A machine-readable inventory of every component and dependency inside a built artifact.