advanced 3 min answer

A gateway gains automatic failover: when the primary model provider returns 5xx or 429 the request is retried against a second provider. On paper, availability rises from 99.5% to better than 99.99%. What has the team given up, and when does that bill arrive?

ai-gatewaysfailoveravailabilityevaluationlock-in
Show the full answer Hide the answer

What is gained, and by how much

If the two providers failed independently, 99.5% each would combine to about 99.9975% — roughly 13 seconds of downtime a month instead of 3.6 hours. That arithmetic assumes independence, and the correlations are real: both providers run in the same few cloud regions, both are capacity-constrained by the same accelerator supply, and the event most likely to rate-limit you — a viral moment in your own product — hits both at once. Treat the gain as "survives one provider's bad hour", not as four extra nines.

What is paid

A second provider is a different behavioural contract, not a second endpoint.

  • Tokenisation differs, so the same prompt is a different token count. A request sitting comfortably inside the context budget on one provider can exceed it on the other, and the failure mode is a hard error at the worst moment.
  • Tool-calling and structured-output surfaces differ in schema shape, in strictness, and in what happens with a malformed call. Code that parses one provider's arguments needs a second parser and a second set of retries.
  • Refusal and safety behaviour differs, so a request that is answered becomes a request that is declined, which your application sees as an empty response rather than an error.
  • The prompt forks. Keeping one prompt that is merely acceptable on both is a quality tax on every request; maintaining two means the evaluation suite doubles, and so does the work for every prompt change afterwards.
  • The prompt cache is per provider and per model. Reading a cached prefix is roughly an order of magnitude cheaper than resending it, so failover loses that discount and raises time-to-first-token exactly when load is highest.
  • Side effects can duplicate. Retrying a tool-calling turn on another provider can re-execute tool calls the first attempt already ran, unless every tool takes an idempotency key.

When the bill arrives

During the first real failover, at peak, with nobody reading output quality. The path is exercised only in incidents, which is precisely when the team is busy and when a silent quality drop is indistinguishable from the incident. Weeks later the complaints arrive with no telemetry to explain them.

How to keep the option open without the silent failure

  1. Make failover a labelled degraded mode, not a transparent swap. The second provider gets its own pinned prompt, a reduced tool surface, and an evaluation run on every change like the primary.
  2. Exercise it continuously. Route 1% of production traffic to the secondary every day and score it with the same judge. A failover path that runs once a quarter is an untested path.
  3. Record the served model on every response, so a quality question three weeks later is answerable.
  4. Shed before you swap. For many products the correct order is: retry with backoff, then queue with a deadline, then degrade the feature, and only then change provider.

When this is the wrong answer

A single product with a 99.5% target and no contractual uptime commitment should not buy this. One provider plus a retry budget, a queue and an honest "try again shortly" is cheaper and more predictable than two behavioural contracts, and the engineering saved pays for the evaluation work that actually moves quality. Multi-provider earns its cost when an uptime clause, a regulator, or a procurement requirement puts a number on the outage.