Unstable was the safe part
OpenTelemetry moved the GenAI conventions into a repository of their own and told instrumentation authors to get the schema URL from there. That repository has no releases, no tags, and eight unreleased breaking changes — so the record of what agents did is being written in a vocabulary with no documented way back.
The argumentOpenTelemetry built the telemetry schema so unstable attribute names could change without orphaning the data already collected, and agent telemetry is now the one body of data being emitted with no version to migrate it from.
Somewhere in a thirty-three kilobyte text file in the OpenTelemetry semantic conventions repository, under the heading 1.27.0, sit two lines that keep a great deal of software working:
gen_ai.usage.prompt_tokens: gen_ai.usage.input_tokens
gen_ai.usage.completion_tokens: gen_ai.usage.output_tokens
They are unremarkable until you consider what happens without them. Transpose the two pairs and every cost report built on that data inverts, silently, while continuing to render. Under 1.30.0 the same file records gen_ai.openai.request.seed becoming gen_ai.request.seed. Under 1.37.0, gen_ai.system becomes gen_ai.provider.name, and three OpenAI-specific attributes shed their gen_ai. prefix entirely.
This file is a telemetry schema, and it is one of the more quietly excellent pieces of design in modern observability. Every batch of telemetry an OpenTelemetry library emits carries a schema URL — a version identifier that also happens to be the address where the schema file can be fetched. A consumer that receives a span written three years ago can read the version it claims, fetch the file, walk the chain of renames between then and now, and present the data in the current vocabulary. The names are allowed to be wrong, because the migration is written down. Unstable was never the dangerous property. Unrecorded was.
In release v1.42.0, the semantic conventions project deprecated every gen_ai.* attribute, metric, event and span it held, and moved them to a repository of their own. The changelog entry is explicit about what instrumentation authors should do next: refer to the new repository for the corresponding schema_url to use.
The new repository's README has a section headed "Schema URL". Its contents are the word TODO.
This is not an oversight hidden in a corner. The repository knows exactly what it intends to do. RELEASING.md describes a dev release channel in full operational detail, down to the shell command that renders the changelog: tags of the form vX.Y.Z-dev, schema URLs under https://opentelemetry.io/schemas/gen-ai-dev/X.Y.Z-dev, releases published as prereleases with the resolved schema attached as an asset. The workflow file that performs it is committed to the repository. model/manifest.yaml already declares schema_url: https://opentelemetry.io/schemas/gen-ai-dev/1.42.0-dev and stability: development.
The repository has no releases. It has no git tags at all. The changelog contains a single heading, "Unreleased", and the marker where generated release notes would begin. The core conventions repository, by contrast, carries twenty-six version tags and has shipped two further releases since the one that handed GenAI away.
So the vocabulary has been in motion ever since, the machinery to record that motion has been designed and written and committed, and it has never produced a release. Meanwhile the data is being emitted. Right now, in production, by agent frameworks and inference clients and MCP servers, into backends with retention policies measured in months.
What is in motion is worth being specific about, because "development status" sounds like a caveat about field names nobody uses yet. Of the seventy-nine gen_ai.* attributes in the registry, not one carries a Stable badge; every single one is marked Development. The agent spans document makes the comparison for you without meaning to: in the attribute table for creating an agent, error.type and server.port, borrowed from the core conventions, wear Stable, and every gen_ai. attribute beside them wears Development.
And the queue of pending breaking changes, each one a file in changelog.d/ waiting for a release that has not come, is eight deep. One renames gen_ai.usage.cache_creation.input_tokens to gen_ai.usage.cache_write.input_tokens. One replaces the gen_ai.client.token.usage histogram, and the gen_ai.token.type attribute it depended on, with separate per-direction histograms. One removes the cache-token breakdown from the invoke_agent span on the entirely sound reasoning that aggregating cache reads across models and inference calls produces a misleading number. One changes gen_ai.request.top_k from a double to an int and moves retrieval's use of it to a new attribute.
Read that list as an operator rather than a spec author. Those are the token and cache attributes — the ones every inference bill, every cost dashboard and every capacity model is built from. They are about to be renamed, retyped and restructured, and when they are, nothing will exist to tell a query engine that last quarter's cache_creation and next quarter's cache_write are the same quantity.
Here the honest complication: a published schema file would not have saved all four. The file format permits a deliberately limited set of transformations — renaming attributes, renaming metrics and events, and splitting a metric. It cannot express a change of value type, so the top_k retyping is beyond it either way. It cannot express a histogram being replaced by three. The schema is a translation layer, not a time machine, and the project's reasons for each of these changes are good ones; this is a convention improving, not a convention failing.
But most of the eight are renames, and renames are precisely what the mechanism handles. More importantly, the schema URL does something that has nothing to do with whether a transformation exists: it states, in the data, which vocabulary that data was written in. That is the part being skipped. A span emitted today says nothing about the convention version it was built against, because there is no version to name. Its origin is a commit on a main branch, and commits are not something a backend can reason about.
The distinction worth holding onto is that stability and versioning are different promises, and OpenTelemetry's own design separates them cleanly. Development status is a statement about the future: these names may change, do not build your business on them. A schema URL is a statement about the past: this data was written in this vocabulary, here is how to read it now. The second is available without the first — that is the whole reason the mechanism exists. The GenAI conventions currently offer neither.
The strongest case against treating this as a problem is a good one, and it is the argument for patience. Premature stabilisation is a genuinely expensive mistake, and a vocabulary frozen too early is one every implementer then works around for a decade. The agent vocabulary is not settled, visibly so — the agent spans document is currently growing conventions for loading a skill, reading a skill resource and executing a command, which is to say the subject itself is still acquiring new parts. Freeze the names in 2026 and you canonise a 2026 theory of what an agent is. And the discipline on display is real: every breaking change has a fragment, the release process is documented to the command line, the version is already chosen. The first dev release could land any week, which would blunt the sharpest form of this complaint and leave a milder one.
But that defence answers an objection nobody is making. Nobody is asking for Stable. The dev channel the project designed for itself would be sufficient — a tag, a resolved schema attached to it, a URL that resolves to a file. The reason it has not happened is not disagreement about names. It is that cutting a release of a convention has no deadline, no incident, and no owner who feels the cost, because the cost is deferred onto someone who will go looking at an agent trace in eighteen months and find a column that stopped existing. The queue of eight is not evidence that the system is working. The queue is the bill.
Which is where this stops being a story about a repository and becomes a question about how we classify the data. Agent telemetry is the only artefact that says what an autonomous system actually did: which tool it invoked, with which arguments, after which plan, at whose expense, and whether a human was in the loop at the moment it mattered. We provision it like monitoring — a retention tier, a sampling rate, a dashboard someone fixes by hand when a field moves. But we are increasingly asked to use it as a record: by a customer disputing an action, by an auditor, by a regulator, by finance reconciling an inference bill, by an incident review reconstructing a decision made eleven months ago.
Monitoring data tolerates an unversioned vocabulary, because nobody queries last March and a broken panel is a morning's work. A record does not tolerate it at all. The uncomfortable arithmetic is that the retention period most organisations have set on agent traces is now longer than the half-life of the words inside them.
Nothing in that arithmetic requires waiting for OpenTelemetry. An architect can stamp a version of their own — the convention commit, the instrumentation version, a small set of attributes named and owned in-house, written alongside the gen_ai.* ones for exactly the handful of facts the organisation will be asked about later. It is unglamorous, it duplicates something a standard should provide, and it is the difference between keeping data and keeping evidence.
Everything above is argued from one party's published record, because it is the party that publishes everything: OpenTelemetry's own repositories, manifests, changelog fragments and specification text, read at the commits cited below.
The schema file in the core repository still carries the line that turns gen_ai.usage.prompt_tokens into gen_ai.usage.input_tokens, a rename made seventeen minor versions ago. It is the most useful sentence anyone has written about agent observability, and it was written by the project that no longer owns the subject. We will keep the traces. What has not been decided is what they will mean.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
- semantic-conventions CHANGELOG, v1.42.0 breaking change — GenAI conventions move to a dedicated repository (read at commit 8b49d42)
- semantic-conventions-genai — model/manifest.yaml, schema_url and stability (read at commit e07f4eb)
- semantic-conventions-genai — RELEASING.md, the dev release channel (read at commit e07f4eb)
- semantic-conventions-genai — releases, "There aren't any releases here"
- semantic-conventions-genai — changelog.d/ unreleased breaking fragments 217, 374, 440, 469 (read at commit e07f4eb)
- semantic-conventions-genai — docs/registry/attributes/gen-ai.md, 79 attributes and their stability badges (read at commit e07f4eb)
- semantic-conventions-genai — docs/gen-ai/gen-ai-agent-spans.md, agent and skill spans and their stability badges (read at commit e07f4eb)
- OpenTelemetry specification — specification/schemas/README.md, Schema URL and schema-aware consumers (read at commit d167c3b)
- OpenTelemetry specification — schemas/file_format_v1.1.0.md, the permitted transformations (read at commit d167c3b)
- semantic-conventions — schemas/1.44.0, the gen_ai renames at 1.27.0, 1.30.0 and 1.37.0 (read at commit 8b49d42)
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.