Goodput
also called Useful Throughput
The rate of completed work that someone is still waiting for, which diverges from throughput under overload and is the only rate that tracks what users receive.
A service reports 12,000 completions per second and the people using it say nothing works. Both statements are true. Nine thousand of those completions are second and third attempts at requests whose first attempt timed out, or they are responses produced after the caller gave up and closed the connection.
Goodput counts only the completions somebody is still waiting for. Throughput counts everything the machine did. Under normal load the two are the same number, which is why nobody instruments the difference until an incident makes it the only question that matters.
Why it matters
Overload does not usually announce itself as a drop in throughput. Every signal an operator trusts keeps looking healthy: CPU is high because the service is working hard, completions per second are high or rising because retries added arrivals, and the error rate can be low because each individual attempt eventually succeeds.
The mechanism that creates the gap is a positive feedback loop. A client that retries twice on timeout converts one user request into three arrivals, so a service slowing down receives more work precisely because it slowed down. Latency rises, more requests pass the timeout, and offered load rises again. Capacity added into this loop is consumed by retries, which is why adding instances during such an incident often changes nothing a user can feel.
Implementation patterns
- Count completions whose deadline had not already passed. Attach an absolute deadline at the edge, check it before returning, and emit two counters rather than one. The ratio between them is the metric.
- Tag retries at the client and split offered load into first attempts and repeats, so amplification is visible as a number rather than inferred from a shape.
- Retry budgets, capping retries at a small fraction of first attempts across the fleet — on the order of 10% — so amplification has a ceiling that does not depend on every client being well behaved.
- Drop expired work at dequeue instead of serving it. A request whose deadline has passed costs the same to serve as one that is still wanted and is worth nothing.
- Alert on the ratio, not the rate. Goodput divided by throughput falling below roughly 0.9 is an overload signal that arrives earlier than latency alerts and far earlier than error-rate alerts.
Industry example
A payments gateway in the mould of Razorpay shows the pattern in its sharpest form. Merchant servers call the gateway synchronously, time out at a few seconds, and retry, because a payment that might not have been taken has to be attempted again. During a festival sale the gateway's completion rate can rise while the number of payments merchants actually observe falls, purely from retries it did not ask for. The remedy follows from the loop rather than from any one company's write-up: admission control at the front door, because capacity added behind it is consumed by the retries.
Failure scenarios
- Retry storms where a brief slowdown is amplified into sustained overload by clients outside your control.
- Queues longer than the deadline, so the service runs flat out and every response is late.
- Hedged requests without a cap, which deliberately duplicate work and silently halve goodput at peak.
- Dashboards measuring throughput only, so an incident is diagnosed as "the service is fine, must be the network" while two thirds of the work is waste.
- Autoscaling on CPU during a retry storm, which adds instances that serve more retries and raises the bill without raising goodput.
Trade-offs
Measuring goodput costs an end-to-end deadline convention and a second set of counters, and the deadline has to be propagated to be checkable. That is platform work, not a per-team change.
Acting on it costs availability on paper. Dropping expired work and capping retries raises the visible error rate while improving what users receive, and that trade has to be agreed before an incident, because during one the error-rate graph is the one on the wall.
When not to use it
For a batch or offline pipeline there is no caller waiting and no deadline, so throughput is the correct and complete metric; inventing a goodput number there adds noise. Goodput earns its place only where a caller gives up, and if nothing times out the two numbers are identical by construction and one of them is overhead.
Interview question
Q: "Your service reports 12,000 requests per second with 1% errors and 80% CPU, and customers report widespread failures. Walk me through what you measure, in order, and what you would change first."
What a strong answer covers: separating first attempts from retries; counting completions whose deadline had not expired; recognising retry amplification as a feedback loop rather than a traffic increase; choosing admission control and retry budgets over capacity; and naming the trade — a higher reported error rate in exchange for work that reaches users.
Quick check
Quiz: Offered load is 12,000 per second and only 4,000 completions per second are still wanted. Name the gap and its two usual causes. Goodput versus throughput; retry amplification and work served after its deadline expired.
Flashcard: Why can throughput rise during an overload while users see a total failure? — Retries add arrivals and expired requests are still served, so the machine completes more work and less of it is wanted.