Throughput
Work completed per unit time, and why it trades against latency.
6 to work through
-
beginner
A team is asked to "make the system faster" for a delivery platform. What are they actually being asked, and why does the distinction change the work?
1 min answer -
beginner
Why does making every request 30% faster often leave requests per second exactly where they were, and what did the optimisation actually buy?
3 min answer -
intermediate
A pipeline moves from writing records one at a time to batches of 500, and throughput rises roughly twentyfold. What has been given up, and at what batch size does the trade stop paying?
2 min answer -
intermediate
Throughput has plateaued and adding instances does not help. What are the candidate bottlenecks and how do you identify which?
2 min answer -
advanced
A capacity test on a live-streaming platform at Twitch scale plateaus at 18000 requests per second however many application instances are added. Fleet CPU is 45%, the database is at 20%, and no queue is growing. What do you check before reporting that ceiling?
3 min answer -
advanced
A recommendation platform must ingest billions of interaction events daily without the ingestion path becoming a bottleneck. Which throughput techniques matter most, and which common approach limits scale?
2 min answer
4 terms in this topic
Collective Communication Cost
The synchronisation term that determines how far a distributed training or inference job can scale, made of a bandwidth component that is nearly flat…
metricGoodput
The rate of completed work that someone is still waiting for, which diverges from throughput under overload and is the only rate that tracks what use…
conceptSaturation Point
The load level beyond which additional demand produces queueing and latency growth rather than additional completed work.
metricThroughput
Work completed per unit time — bounded by the system's narrowest resource, and traded against latency.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.