advanced 3 min answer

You are asked to log every prompt and response for an assistant handling 2 million interactions a day, so quality regressions are diagnosable. Roughly how much data does that commit you to per year, and does the number change the design?

observabilityloggingretentioncostprivacy
Show the full answer Hide the answer

The assumptions, stated

  • 2 million interactions per day.
  • Average request: system prompt 600 tokens, retrieved context 2000, conversation history 400, so about 3000 input tokens.
  • Average response: 400 output tokens.
  • Roughly 4 bytes per token of UTF-8 text for English prose.
  • Structured metadata per interaction: model snapshot, prompt version, retrieval document ids and scores, per-stage latencies, token counts, guardrail verdicts. About 1 KB.

The arithmetic

3400 tokens at 4 bytes is about 13.6 KB of text, plus 1 KB of metadata, so call it 15 KB per interaction.

2 million x 15 KB = 30 GB per day, which is about 11 TB per year raw.

Text compresses well, and telemetry of this shape stored columnar compresses better because the repeated system prompt and the metadata columns are highly redundant. At 6x, the stored volume is roughly 1.8 TB per year. On object storage at a few cents per GB-month, the storage bill is tens of dollars a month, not thousands.

The number, with its range

11 TB raw and roughly 1.5 to 2.5 TB stored per year, and storage cost is not the constraint. The range is dominated by one assumption: retrieved context size. If retrieval grows from 2000 to 12000 tokens, which is an ordinary product decision, the raw figure goes to about 38 TB per year. Context length, not interaction count, is the term that decides this estimate, and it is the term nobody checks before signing off the design.

What the number rules in and rules out

It rules in full retention, and it rules out three designs people reach for anyway:

  • Logging synchronously on the request path. 30 GB a day of writes coupled to user-facing latency means the logging store's availability becomes the assistant's availability. Write to a local buffer or a queue, and accept loss under extreme load rather than accepting a new hard dependency.
  • Keeping it in the primary transactional database. 11 TB of append-only text in Postgres alongside the application's own tables makes every backup, restore and major-version upgrade worse. Columnar object storage with a query engine is the shape that fits.
  • Indefinite full-fidelity retention. Not because of bytes, but because prompts contain whatever users typed, which at this volume certainly includes personal data. Retention becomes a legal position: deletion on request must reach this store, residency rules may forbid a single global bucket, and every engineer with query access can read customer conversations.

The design that follows

Full fidelity for 30 days, then sampled. Keep 100 percent of interactions that were flagged, thumbed down, abstained, hit a guardrail or exceeded a latency threshold, plus 1 to 5 percent of the rest, for 12 months. That preserves the rare cases regression analysis needs while cutting the long-tail volume by more than an order of magnitude. Record the sample rate on every row so counts can be reconstructed by weighting; without it, every historical count silently understates volume.

When this is over-engineering

Below roughly 50000 interactions a day, the whole discussion is moot: keep everything for 90 days, in whatever store you already have, and spend the saved effort on the labelled evaluation set that will actually tell you whether quality moved. The sampling machinery earns its cost at volumes where the privacy surface, not the storage bill, is what you are managing.