AI Agent Orchestration Platform  ·  View 26 of 32  ·  5 · Operations

Cost Management and Attribution

How every token and tool call is metered, attributed, capped and then made cheaper.

Editable source SVG draw.io All views
Meter Token Counting gateway policy Tool Call Units priced per API Infrastructure Cost Cost Management Attribute Execution Tagging tenant · project Rollup Job hourly · ADX Shared Cost Split by token share Enforce Pre-Call Estimate reserve tokens Soft Limit warn at 80 percent Hard Limit block at ceiling Optimise Cost-Aware Routing cheapest passing Cache Hit Rate semantic + exact PTU Utilisation reserve vs burst Context Compression tokens per run Report Cost Dashboard per tenant · agent Forecast run-rate to month end Chargeback Export monthly · CSV degrade first raise threshold avoided tokens Cost Management — Meter, Attribute, Enforce, Optimise Interface / broker Security / platform Application we own Decision point Data store failure / alternate synchronous A breached budget degrades the route before it blocks the run. Blocking is the last step, not the first. v 1.0 · owner FinOps and Architecture · date 2026-08

Decisions

  • Metering happens at the gateway, not in the agent, so an agent cannot under-report its own consumption
  • Budget is reserved at admission and reconciled at completion; a run cannot start that it cannot afford to finish
  • A breached budget degrades the route to a cheaper model before it blocks the run — blocking is the last step, not the first

Attribution

  • Per execution, per agent, per workflow, per project, per tenant, and a shared-infrastructure split by token share
  • Alerts at 80 percent of budget, hard stop at the ceiling, with a monthly forecast from the current run rate
  • Chargeback export monthly per tenant, reconciled against Azure Cost Management for the infrastructure component

The levers, in order of value

  • Context compression — tokens per run is the single largest cost driver and the one the platform controls directly
  • Semantic and exact-match caching, measured as a hit rate per agent
  • Cost-aware routing to the cheapest model that passes the agent's quality bar, and PTU utilisation kept high enough to justify the reservation