You cannot bill a histogram
OpenTelemetry's GenAI conventions describe token usage fourteen ways on a span and one way in a metric. The fourteen can be priced and get sampled; the one is kept in full and cannot tell a cached token from a full-price one. Cost attribution for agents is being built on the gap between them.
The only complete token signal in the GenAI conventions cannot tell a cached token from a full-price one, and everything that can price a call rides on spans tracing exists to discard, so AI cost attribution runs on a measurement nobody guarantees is complete.