advanced 3 min answer

Zoom publicly describes the architecture behind its AI Companion as a federated approach - its own models used alongside third-party frontier models, with work routed by task rather than every request going to a single provider. What problem does that structure solve that a single-provider design does not, and where would copying it be a mistake?

model-routingmulti-providercostevaluationzoom
Show the full answer Hide the answer

The situation that shapes the choice

A communications platform puts AI on a path that every meeting crosses: transcription, summarisation, action items, chat replies. Two properties follow from that and they drive the whole design.

Volume is enormous and the per-interaction value is low. A meeting summary is useful and nobody will pay a large amount for each one, so cost per interaction is a binding product constraint rather than a finance concern. And the task mix is wildly unequal in difficulty: summarising a transcript against a fixed template is close to a solved problem for a small model, while multi-step reasoning over a user's calendar, documents and history is not.

What the structure buys

  • The cheap majority runs cheaply. If most traffic is the easy shape, serving it with a small owned model and escalating only the hard minority changes the unit economics of the whole feature. Routing pays exactly in proportion to how skewed the difficulty distribution is.
  • No single provider is a single point of failure. Provider rate limits, regional capacity shortfalls and deprecation timetables stop being existential. A product embedded in a meeting cannot show an error because one vendor is throttling.
  • Procurement answers get easier. Enterprise buyers ask which model processed their data and where. A design that already routes per task can answer per task, and can offer a configuration that excludes a given provider.

What it costs

  • Evaluation multiplies. Quality must be established for every task on every candidate model, and re-established whenever any of them changes. N models times M tasks of evaluation is the real bill, and it recurs.
  • Prompts are not portable. A prompt tuned for one model's instruction-following behaviour underperforms on another. Every routed task needs per-model prompt variants under version control, which is the prompt-management problem multiplied by the fleet.
  • The router is on every request path. It becomes a component with its own latency, its own failure modes and its own capacity for a bad configuration to send all traffic to the expensive model at 03:00.
  • Running your own model is an organisation, not a project: training data, evaluation, serving capacity, on-call.

When not to copy this

At 50000 interactions a day, this design loses money. The routing layer, the per-model evaluation harness and the duplicated prompt variants cost more engineering than the token savings are worth, and they add failure modes to a product that has not yet proven anyone wants it. One provider, one dated model snapshot, a documented fallback and a measured evaluation set is the correct architecture until model spend is a top-three line item.

The second mistake is copying the structure without the measurement that justifies it. Routing only pays when task difficulty is genuinely separable and you can tell, per request class, which model is sufficient. Without that evidence the router becomes a random assignment of users to quality levels, and the complaints will not correlate with anything you are logging.

The decision rule

Introduce routing when one task class both dominates volume and is demonstrably easy, measured on your own labelled set, and when the saving at current volume pays for the evaluation harness within a couple of quarters. Keep the abstraction thin: a per-task model binding in configuration, not a bespoke routing service, until the bindings themselves need logic.