Term Kind Topic What it is
Interaction to Next Paint INP, Responsiveness Metric metric Real User Monitoring A field metric reporting one of the slowest complete interactions on a page - input delay plus handler processing plus the wait for the next paint - which replaced First Input Delay as a Core Web Vital on 12 M…
Internal Consumer SLO metric Platform SLOs A reliability commitment made to teams who cannot switch supplier, which is why it must be measured from their side rather than from the platform's.
Keyed State Size metric Stateful Stream Processing The total state a job holds per key across all keys, which governs memory, checkpoint duration and recovery time.
Last Safe Start Date Programme Latest Start, Deadline Backstop Date metric Negotiation The externally imposed deadline minus the programme's duration and contingency - the date an architect computes and owns, which competes for funding in a way the distant vendor deadline does not.
Latency metric Latency How long one operation takes — a distribution, never a number, and dominated by its tail in any system with fan-out.
Latency Revenue Elasticity metric Performance vs Cost The measured change in a business outcome per unit of added or removed latency, which turns "faster is better" into an amount of money and tells you when to stop buying milliseconds.
Lead Time for Changes metric Flow Metrics The elapsed time from code committed to code running in production.
Lead Time to Value metric Time to Market The elapsed time from identifying an opportunity to delivering measurable benefit, which is usually dominated by waiting rather than by building.
Lineage Coverage Harvest Coverage, Lineage Completeness Ratio metric Data Catalog The share of known data assets whose upstream and downstream edges were successfully derived on the latest harvest, which is the only signal that separates an asset with no dependencies from one the parser cou…
Little's Law in Practice metric Little's Law L = λW — concurrency equals arrival rate times latency — and the reason a slowdown becomes an outage.
Long Task Main-Thread Blocking Task, Long Animation Frame metric Web Performance Budgets Any single piece of work holding the browser's main thread for more than 50 ms, during which no input is handled and nothing is painted - the unit in which an unresponsive page is actually measured.
Match Confidence Band Match Threshold Band, Review Band metric Master Data Management The score range between automatic merge and automatic reject in record matching, which is where a consolidation programme's permanent staffing cost lives and where its worst errors are prevented.
Mutation Score metric Mutation Testing The proportion of deliberately introduced faults that the test suite detects — a measure of whether tests would notice a defect, unlike coverage.
Overload Recovery Time Overload Settling Time metric Stress Testing Wall-clock time from the moment load returns to normal until the system is fully healthy again, which is set by spare capacity rather than by service speed.
Page Budget Pages Per Shift, Alert Budget metric On-Call An explicit ceiling on how many pages a shift may generate, treated as a limit the team manages against rather than as an outcome it observes.
Path Exit Rate Golden Path Abandonment Rate, Step Abandonment Rate metric Paved Road & Golden Path The share of journeys along a supported path that stop at each step, which tells a platform team where teams are forced off the road rather than merely how many ended up off it.
Peak Multiple Headroom Multiple, Peak-Over-Peak Factor metric Estimation The factor by which a system is sized above its last observed peak - a fast planning heuristic that silently assumes the population generating load is the population you served before.
Per-Device Series Budget Fleet Series Count, Device-Tag Cardinality Limit metric Device Telemetry at Scale The number of distinct metric time series a fleet creates - series per device multiplied by device count - which decides observability cost long before data volume does, and which determines whether per-device…
Percentile Latency metric Latency Latency expressed as the value below which a given proportion of requests fall, used because averages conceal the behaviour that users notice.
Physical Latency Floor Speed-of-Light Budget, Propagation Floor metric First-Principles Reasoning The minimum achievable round-trip time between two locations set by propagation in fibre - roughly 1 ms per 100 km each way - which no provider, protocol or cache can improve on for a request that must reach t…
Pipeline Change Coverage Change Denominator, Measured Change Share metric Flow Metrics The share of changes reaching production that the delivery pipeline actually measures - the denominator that decides whether elite delivery metrics describe the system or only the part of it that happens to be…
Pipeline Feedback Latency Commit-to-Verdict Time, CI Feedback Time metric CI/CD The time from pushing a change to knowing whether it passed, which sets the batch size engineers choose and therefore the size of every review, deploy and rollback.
Platform Adoption metric Platform Adoption The proportion of eligible teams and workloads actually using the platform, treated as the platform's primary product metric.
Platform SLO metric Platform SLOs A published reliability and performance objective the platform commits to for the teams that depend on it.
Projection Rebuild Budget Read Model Rebuild Time, Rebuild Service Level Objective metric CQRS The measured wall-clock time to rebuild a read model from scratch, treated as a service level objective, because it bounds how quickly a projection defect can be corrected and therefore whether CQRS is reversi…
Quality Dimension Threshold metric Data Quality Dimensions The stated numeric level at which a dataset is fit for its purpose on a given quality dimension, plus what happens when it is not met.
Quality Proxy Metric Implicit Quality Signal, Behavioural Quality Indicator metric AI Observability A continuously available production signal that moves when answer quality moves - abstention rate, regeneration rate, escalation rate - used to page someone in the hours before a labelled evaluation could ever run.
Queue Time Attribution metric Flow Metrics Splitting elapsed lead time into work and waiting, and naming which queue each wait sat in.
Radio Wake Budget Transactions Per Day Budget, Wake Count Budget metric Constrained Protocols The number of radio wakes a device can afford per day at its target battery life - the planning number that decides protocol and batching choices, because energy is spent per transmission rather than per byte.
Recovery Point Objective RPO metric RTO & RPO The maximum acceptable data loss measured as a duration, determining the replication and backup strategy.
Recovery Time Objective RTO metric RTO & RPO The maximum acceptable duration between a failure and restored service, agreed with the business rather than chosen by engineering.
Reliability Headroom metric Capacity Planning The gap between provisioned capacity and the load that would be carried after the largest planned failure, measured under peak conditions.
Remote Resolvability Issues Fixable Without a Visit, Field Visit Rate metric Fleet Management The proportion of device issues that can be diagnosed and resolved without a physical visit - a design-time property that determines the operating cost of a fleet for its entire life.
Replication Lag metric Replication How far behind a replica is, measured in time or in log position — the quantity that determines how stale a replica read can be.
Request-Based and Window-Based SLI metric SLI, SLO & SLA Two ways of computing the same reliability target — counting good events, or counting good time windows — which produce materially different numbers.
Retention Cost metric Streaming Cost The storage bill for keeping a log replayable, which is set by retention multiplied by throughput multiplied by the replication factor.
Review Latency metric Code Review The time between a change being ready for review and the review happening, which is usually the largest component of lead time.
Reviewer Throughput Ceiling Oversight Capacity Limit, Review Time Budget metric Human-in-the-Loop Design The number of items a human reviewer can genuinely assess per hour, which bounds what any human-in-the-loop control can actually deliver regardless of what the design document claims.
Round-Trip Time RTT, Latency Floor metric Network Performance The time for a packet to travel to a destination and back — a physical floor that no application optimisation can reduce.
RTO and RPO metric RTO & RPO How long recovery may take, and how much data may be lost — the two numbers from which every disaster recovery design follows.
RTO and RPO Recovery Time Objective, Recovery Point Objective metric Reliability & Resilience How long recovery may take (RTO) and how much data may be lost (RPO), the two numbers that determine the cost of a resilience design.
Sampling Risk Detection Probability, Audit Sampling Power metric Continuous Controls Monitoring The probability that a sample-based control test misses a violation that exists in the population - which for small samples and rare violations is most of the time.
Scanned Bytes metric Analytics Cost Control The volume a query reads, which is what most analytical engines bill for and what almost every optimisation ultimately reduces.
Service Level Indicator SLI metric Reliability & Resilience The actual measurement of a service's behaviour that an objective is set against — a ratio of good events to valid events.
Service Level Objective SLO metric Reliability & Resilience An internal target for a service level indicator, set below the level at which users notice, and used to decide whether to ship or to stabilise.
SLI, SLO and SLA metric SLI, SLO & SLA The measurement, the internal target, and the external contract — with the error budget as the mechanism that makes the target consequential.
Slip Cost per Week Weekly Delay Cost, Delay Cost per Week metric Time to Market The money a launch loses for each week it is late - and whether that money is deferred or destroyed - which is the only form in which time to market can be traded against design quality.
Supported Path Maintenance Load Golden Path Upkeep Cost, Path Matrix Load metric Paved Road & Golden Path The recurring engineering cost of keeping every supported path working against every change in the platform beneath it, which is what actually caps how many paths a platform can offer.
Tail Latency p99, p999 metric Performance & Capacity The latency experienced by the slowest small percentage of requests, which is what users and dependent services actually feel.
Telemetry Pipeline Lag Ingestion Delay, Time to Queryable, Observability Freshness metric Observability The delay between an event being emitted and being queryable, which bounds how quickly any decision made from telemetry can respond to reality.
Telemetry Spend Ratio metric Observability Cost Observability cost as a proportion of the infrastructure it observes, used as a tripwire for a category that grows silently.
Third-Party Weight Budget Tag Budget, Non-Application Script Allowance metric Web Performance Budgets An enforced allowance for scripts the application build does not produce - count, bytes and main-thread time - with a named owner and a technical enforcement point, because these are the scripts a bundle-size …
Throughput metric Throughput Work completed per unit time — bounded by the system's narrowest resource, and traded against latency.
Time to Detect MTTD metric Incident Management The interval between a problem beginning and someone knowing about it, which is often the largest and most reducible component of total incident duration.
Time to Interactive Gap metric Hydration Cost The interval between a page appearing complete and actually responding to input, during which the user's taps do nothing.
Total Cost of Ownership TCO metric Cost & FinOps The full lifetime cost of a capability, including the people, operations, upgrades and exit that a licence comparison leaves out.
Unit Cost Ceiling Cost per Transaction Limit, Infrastructure Cost Budget per Order metric Business Understanding The infrastructure cost per transaction that the business model can bear, derived from contribution margin and fixed before design, which decides how much work a request is allowed to do.
Unit Economics metric Unit Economics Infrastructure cost expressed per unit of business value delivered, which reveals efficiency trends that absolute spend cannot.
Unit Economics of a System metric Unit Economics The cost to serve one unit of business value — a customer, an order, a request — which determines whether a system's economics improve or worsen with scale.
Utilisation Break-Even Break-Even Utilisation, Execution Model Crossover metric Serverless vs Containers The sustained busy fraction at which always-on compute becomes cheaper than per-invocation billing, which turns the serverless-or-containers argument into arithmetic instead of preference.