Source Load Budget
also called Upstream Read Quota, Analytical Load Cap
An explicit cap on the concurrency, rows and duration that analytical reads may consume on an operational source - so a dashboard query cannot become a production incident.
An analyst opens a dashboard built on a federation layer. One tile joins two tables on an operational replica with no selective filter, the replica spends four minutes on the scan, replication lag climbs, and within ten minutes the application is serving stale order status to customers. Nothing was misconfigured and nothing errored. The analytical path had no limit on what it could take from the operational path.
A source load budget makes that limit explicit: how many concurrent analytical connections a source accepts, how many rows or bytes one query may read, how long a statement may run, and which replica it runs against.
Why it matters
Reading in place is attractive for good reasons - no copy to secure, no second retention clock, no staleness. But the two query shapes are incompatible: one wants a few rows by key in single-digit milliseconds, the other wants a hundred million rows and will take whatever the machine has. On one database the second starves the first through buffer-pool eviction, snapshots that block vacuum or log truncation, and replication lag.
The asymmetry is what makes the budget worth setting in advance: the saving is a pipeline nobody had to build, and the exposure is customer-visible failure in the system that takes money.
Implementation patterns
- A dedicated replica for analytics, with the federation layer holding credentials to that host only. The cheapest and most effective control, because it caps the damage instead of relying on good queries.
- A connection ceiling per source - a handful, not dozens - so one slow source cannot consume every worker thread and stall queries that never touch it.
- Statement timeouts set at the source, not only in the client, because a disconnecting client leaves the server working, plus a row or byte cap per scan so the query fails rather than degrades.
- Pushdown verification in review: read the plan and confirm filters reach the source. A predicate applied after the fetch means the budget is spent on rows nobody wanted.
- Scheduled extraction for anything read repeatedly, and log-based change capture where the source must feed analytics continuously - the log is the one read path whose cost scales with change volume rather than with table size.
Industry example
Notion has described its first analytics pipeline as roughly 480 hourly connectors pulling from a Postgres estate of 480 logical shards into a managed warehouse, and the overhead of monitoring those connectors and re-syncing them whenever the estate was resharded or upgraded as the reason it did not last. From 2022 it moved to change capture through Debezium into Kafka, writing updates into Apache Hudi tables on S3 (Notion engineering blog, 2024).
Read as a budget decision: the pull path's cost scaled with the number of shards and with every refresh, while the log path's cost scales with the number of changes - for an update-heavy workload, a different order of magnitude of load on the source.
Failure scenarios
- Replica lag that reaches the application, so an analytical query becomes a customer-facing freshness bug.
- A federated query holding connections in four sources for minutes, exhausting the layer's worker pool, so queries against unrelated sources queue behind it.
- A long analytical snapshot preventing log truncation, filling the source's disk - the analytics query becomes a database outage.
- A pushdown regression after a connector upgrade: the query that read 2 MB last month now fetches the whole table.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Tight budget at the source | Queries fail fast and visibly | Some legitimate questions are blocked |
| Generous budget | Exploratory work is easy | The first unbounded query is an incident in production |
| Copy to an analytical store | No source load at all | A pipeline, a second copy and staleness |
When not to use it
When the data has already been copied. A budget protects a shared source; a warehouse table you own needs cost control instead. It is also unnecessary when the source is a read store built for analytics, and for a lookup table of a few thousand rows, where federating with no ceiling costs nothing.
Interview question
Q: "Your BI tool queries an operational Postgres replica through a federation layer. It has been fine for a year. What would you put in place today, and how would you justify the week when nothing is broken?"
What a strong answer covers: that the risk is latent and grows with query authors rather than with data volume; the four controls (dedicated replica, connection ceiling, statement timeout, row cap) and which is cheapest; replication lag during the busiest analytical hour as the measurement that makes the case; and the rule that converts a repeatedly federated query into a scheduled materialisation.
Quick check
Quiz: A federated query raised replica lag to 40 minutes and stale order status reached customers. Which single control would most cheaply have prevented it? A dedicated analytics replica, so the worst outcome is a slow dashboard rather than a stale application.
Flashcard: What does a source load budget cap, and what does hitting it force? — Concurrency, rows or bytes and statement duration per source; hitting it forces a choice between rewriting the query and materialising the data on a schedule.