Concurrency
Operations in flight, and the limits that are the real capacity ceiling.
4 to work through
-
beginner Multiple choice
A service handles 100 requests per second with a pool of 10 worker threads. The team raises the pool to 100 threads. Throughput stays at roughly 100 rps and p99 latency gets much worse. What is the primary reason?
2 min answer -
advanced Multiple choice
A platform must handle millions of concurrent connections per process cluster where most connections are idle most of the time. Which concurrency model fits, and what does the wrong choice cost?
2 min answer -
advanced Multiple choice
A service holds a concurrency limit of 60 with a 50 ms mean service time and serves clients that give up after 2 seconds. Offered load triples. Where should the excess requests wait?
3 min answer -
advanced Multiple choice
Why is limiting concurrency more effective than limiting request rate when protecting a service?
2 min answer
3 terms in this topic
Concurrency
How many operations are in flight at once — the quantity that actually saturates a system, and the one most worth limiting explicitly.
conceptDeadline-Derived Queue Depth
The rule that a request queue's maximum depth comes from the client's deadline and the service's completion rate, so the queue can never hold work th…
conceptParallelism and Concurrency
The distinction between structuring work so it can make progress independently and actually executing work simultaneously on multiple cores.
Neighbouring topics
Performance & Capacity
General material on performance and capacity engineering.
Latency
Distributions rather than averages, and the floors physics imposes.
Throughput
Work completed per unit time, and why it trades against latency.
Queueing Theory
Why latency explodes as utilisation approaches capacity.
Little's Law
L = λW, and the pool sizes it computes directly.
Bottleneck Analysis
Finding the constraint, and expecting a second one behind it.
Tail Latency
p99 behaviour, amplification across fan-out, and hedged requests.
Load Testing
Realistic data, realistic mix, and a ramp rather than a step.
Stress Testing
Pushing past target to learn what breaks first and how it fails.
Soak Testing
Long runs that surface leaks and slow degradation.
Capacity Modelling
Arithmetic before load tests, and headroom for failure as well as peak.
Horizontal vs Vertical Scaling
Scale out for stateless, scale up first for stateful.
Caching for Performance
Layer choice, hit ratio as a first-class metric, and cold-cache recovery.
Database Performance
Plans, indexes, contention and the pool in front of the database.
Connection Pooling
The most common hidden ceiling, and the metric nobody collects.
Network Performance Tuning
Keep-alive, compression, payload size and round-trip elimination.
Performance Budgets
Targets enforced in CI so regressions fail the build.
Profiling & Optimisation
Measuring before optimising, and optimising the dominant term.
Peak Event Readiness
Freeze, pre-scale, shed order, warm caches and rehearse.