Pass-Through Cost of Goods Sold
also called Third-Party Unit Cost, Supplier Pass-Through Term
The part of the cost of serving one unit that is paid to a third party per transaction, which infrastructure optimisation cannot reduce and which eventually decides whether unit cost keeps falling.
A platform's cost per request falls steeply for two years and then stops, while every efficiency programme keeps reporting wins. Compute per request is down 60%, storage per request is down 40%, and the number the business cares about has not moved since the last funding round. Nothing is broken. The costs the team can see stopped being the costs that matter.
For a messaging platform most of what leaves the business per message goes to carriers. For a travel search product it goes to supplier APIs billed per call. For an assistant it goes to model tokens. For a design tool it goes to content royalties per export. These are pass-through terms: incurred per unit of business value, paid to somebody else, and completely invisible to the infrastructure telemetry that engineering built.
Why it matters
The composition of unit cost determines where engineering effort pays. If the pass-through term is 80% of cost of goods sold, halving the infrastructure term removes 10% of unit cost, which lands as roughly one to two points of gross margin. That is worth having, and it is not where a year of engineering should go.
It also changes what the cost curve means. Infrastructure cost per unit generally falls with scale because fixed floors amortise; pass-through cost per unit generally does not, because it is a price times a quantity. A platform that models its future margin on the early curve, when infrastructure dominated, will commit to prices it cannot serve profitably at ten times the volume. That mistake is made in pricing and sales, months before engineering notices.
Implementation patterns
- Decompose unit cost into two lines before setting any target: controllable and pass-through. Publish both. The decomposition usually takes an afternoon and changes the roadmap.
- Meter the supplier charge at the point of use, not from the monthly invoice. Attach the expected supplier cost to the transaction record as it happens, then reconcile against the invoice. A daily feedback loop makes routing changes measurable; a billing-cycle loop does not.
- Make supplier choice data rather than code: a table of destination, supplier, price, quality and effective date, with one adapter interface per supplier. A price change should be a configuration change.
- Route on price and quality together. Least-cost routing that ignores delivery quality moves cost onto retries and support.
- Negotiate against measured volume. Tiered supplier contracts reward committed volume, so the forecast is itself a cost lever, and it is only as good as the per-route metering.
Industry example
Communications platforms of the kind Twilio operates are the clearest case: the published shape of that business is a per-message price to the customer against a per-message carrier charge, so gross margin is set by routing and contracts far more than by server efficiency. The same structure appeared across the industry from 2023 onward in products built on third-party model APIs, where token cost per answer became the dominant unit cost within a quarter of launch and teams discovered that prompt and context design, not cluster efficiency, was the margin lever.
Failure scenarios
- Unit cost flattens and nobody can explain it, because dashboards cover only infrastructure. The symptom is a healthy efficiency report next to a flat margin.
- Pricing is set against the early curve. Contracts signed when infrastructure dominated become unprofitable at scale, and they cannot be repriced for a year.
- A supplier raises prices and the change needs a deployment, so the response takes six weeks instead of an hour.
- Least-cost routing degrades quality, and the saving reappears as retries, refunds and support load that nobody attributes back to the routing change.
- Invoice-only reconciliation, so a routing error is discovered 40 days later.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Route on price per transaction | Directly moves the dominant cost term | Routing complexity and a quality regression risk that needs its own monitoring |
| Single supplier at committed volume | Best tier pricing and simple operations | Concentration risk and no bargaining position at renewal |
| Meter per transaction | Daily feedback and defensible reconciliation | A durable metering pipeline with correctness obligations |
When not to use it
When cost of goods sold is almost entirely your own compute, the split is trivial and the infrastructure target is the right one: an inference platform running its own models, a search backend, a transcoding service. Do the decomposition once anyway, because it tells you which kind of business you are in, and the answer changes the moment the product adds a third-party component. Do not build per-transaction supplier metering when there is one supplier on a flat annual fee, where the pipeline costs more than it can ever reveal.
Interview question
Q: Your platform's cost per request has been flat for three quarters while infrastructure efficiency improved 30%. What do you look at, and what would you change?
What a strong answer covers: decomposing unit cost into controllable and pass-through terms; recognising that the pass-through term does not amortise with scale; metering supplier cost at the point of use rather than from invoices; treating routing and contracts as architecture; and naming the second-order risk that least-cost routing shifts cost into quality.
Quick check
Quiz: A platform's cost of goods sold is 80% supplier fees. The team halves infrastructure cost per unit. Roughly what happens to gross margin? It removes about 10% of cost of goods sold, on the order of one to two points of margin.
Flashcard: Unit cost stopped falling although efficiency improved. What is missing? The pass-through term, paid per transaction to a third party. It does not amortise with scale, and it is controlled by routing and contracts rather than by clusters.