advanced 2 min answer

At 09:40 a release begins logging the whole inbound request as nested JSON, including a customer-supplied metadata map. At 10:05 the log store starts rejecting documents for that index with an illegal-argument error. By 10:20 no service writing to that index can log at all, including services that changed nothing. What happened, and which design decision allowed it?

structured-loggingschemacardinalityblast-radiuselasticsearch
Show the full answer Hide the answer

The trigger

Dynamic mapping creates an index field for every distinct JSON key it has not seen. The customer metadata map has no fixed key set, so a few hundred customers each sending their own keys add hundreds of fields within minutes. Elasticsearch caps this with index.mapping.total_fields.limit, which defaults to 1000, and the 25 minutes between 09:40 and 10:05 is simply how long it took to fill.

Past the limit the store rejects the entire indexing request, not the offending field. Log shippers write in bulk, so well-formed documents batched next to a rejected one die with it.

Why it propagated beyond the guilty service

The mapping is a property of the index, not of the writer. One service's schema change is therefore a write outage for every service sharing that index, and none of them can fix it. Blast radius here is set by index layout, which nobody thought of as a coupling decision.

Why detection lagged

Every service's own view was healthy: the write to stdout succeeded, so no error surfaced in the application. The rejection is counted in the shipper or the store, and those counters are rarely paged on. The symptom reached humans as empty dashboards, which reads like a query problem, so the first stretch of investigation went to the dashboard tool.

The structural fix versus the tempting local one

The tempting fix is to raise the limit to 10000. It buys days and pays in cluster state: every mapped field costs memory in the mapping that is distributed to every node, mapping updates get slower, and searches on a wide index degrade. The limit exists because the failure it prevents is worse than the one it causes.

The structural fix is that user-controlled keys must never become schema. Put the map into a single field with a type that stores arbitrary JSON without mapping subfields — flattened is the usual choice — or serialise it to one string. Set dynamic: strict on the index so an unplanned field fails a release in staging instead of the cluster in production. Give each service its own index or data stream so a schema mistake has a one-service blast radius, and alert on rejected-document counts per index with a threshold of zero, because any rejection is a defect.

The general lesson

Every place where untrusted input can create a new named dimension is an unbounded-cardinality hole whose blast radius is larger than the service that opened it. Log fields, metric labels and span attribute keys are the same hazard in three clothes: the name is schema, the value is data, and only values may come from users.

When not to lock the schema down

A small internal system with a fixed schema and a known set of writers benefits from dynamic mapping, and dynamic: strict on a fast-moving internal app creates a release-blocking chore for a risk that is not there. The flip condition is whether any field name can originate outside your own code. Once one can, strictness costs a few pull requests and prevents an outage that takes a cluster-wide mapping change to undo.