advanced 3 min answer

A partial-update RPC takes the message and writes every field it contains. After a client release, support finds accounts whose discount percentage has become 0 and whose nickname has become empty, with no failed calls and no bad data in the client's logs. The protobuf definition uses plain proto3 scalars. Explain the failure.

protobuffield-presencepartial-updatefieldmasksilent-data-loss
Show the full answer Hide the answer

The trigger

A client sent an update message that set only email. The server deserialised it, read discount_percent as 0 and nickname as "", and faithfully wrote both.

Why it propagated

A plain proto3 scalar field has no presence information on the wire. The encoder omits a field whose value equals the type default, and the decoder cannot distinguish "omitted" from "explicitly set to the default". 0, "" and false are indistinguishable from absent. There is no sentinel, no hasField, and nothing for the server to check.

So the server's rule, "write every field in the message", is not implementable on this message type. It reads defaults for fields the client never mentioned and treats them as instructions. The two values that got destroyed are exactly the ones whose legitimate values collide with the type default, which is why the bug looks selective and random rather than total: accounts with a non-zero discount lost it, accounts already at zero were unaffected.

Detection lagged for the same reason. There is no error path: the RPC returns OK, the write succeeds, and the audit log records a valid update. The only signal is a distribution shift — a spike in the count of accounts at exactly 0 — which nobody alerts on.

The structural fix versus the tempting local fix

The tempting fix is for the client to always send every field. That works until the next client, the next language binding, or the next field, and it makes every caller responsible for the server's correctness.

Two real fixes, in order of preference:

  1. Make presence explicit. Mark the fields optional, which proto3 has supported for scalars since protobuf release 3.15 (February 2021). The generated code then carries has_discount_percent(), and the server can apply only what was actually set. The wire change is compatible: the field number and type do not move, and old readers are unaffected.
  2. Carry an explicit update mask. A google.protobuf.FieldMask naming the paths to update makes the client's intent part of the request rather than an inference from the payload. This is the right answer when the API must also express "clear this field to its default", which presence alone handles but a mask documents, and when nested messages are in play.

Then add the guard that would have caught it: reject an update whose mask is empty or whose message sets no field with presence, rather than treating it as "update everything to default".

The general lesson

An absent value and a default value are different facts, and a format that conflates them cannot express a partial update. The same trap appears in JSON APIs as the null ambiguity in merge patch, and in column-oriented pipelines where a missing column becomes a zero. Any time your API has a "leave this alone" semantic, check that the encoding can represent it before writing the handler.

When not to make presence explicit

If the operation is genuinely a full replacement, say so: name it Replace, document that omitted fields are cleared, and require the client to send the complete object. Full replacement is easier to reason about than a mask and strictly safer than an ambiguous partial update, and for a small resource it is the better design. The cost is bandwidth and a lost-update race that conditional requests, not field masks, solve. What is indefensible is the middle position this API is in: partial-update semantics on an encoding that cannot say which fields were sent.