A building-controls platform stores a desired state per device and reconciles it on every connection. An operator sets a fan setpoint that one hardware revision cannot accept, and applies it in bulk to 40,000 units. The platform compares reported state with desired state to decide whether to act. What happens over the next day, and which single change stops it?
Show the full answer Hide the answer
Minute by minute, what happens
The first affected device connects, receives the desired setpoint, refuses or clamps it, and reports the value it actually holds. The platform compares reported with desired, finds them different, and issues the write again. Equality of values is the only termination condition in this design, and it can never be reached, so the loop has no exit.
At a fifteen-minute reconcile cadence, each affected device exchanges two messages per cycle. Across 40,000 units that is 40,000 × 2 × 96, about 7.7 million messages a day whose only effect is to confirm that nothing changed. Each one also writes a state-change event into telemetry, so the storage bill grows at the same rate.
Where it amplifies
The operator console is the worse casualty. Every affected device shows "pending" forever, so the queue of genuinely pending changes is now buried in 40,000 permanent entries, and the next real configuration push cannot be verified by anyone. The twin has stopped being a model of the fleet and become a record of an argument.
For battery or cellular devices the same loop lands on the metered link: two extra messages every fifteen minutes is roughly 5,800 messages per device-month that nobody asked for.
What stops it, and why the other options fail
- Retry on every connection until accepted is the current design, restated. The device has already answered; repeating the question is not a strategy.
- Clamping silently makes reported equal desired and ends the loop, which is worse. The platform now reports that 40,000 devices hold a setpoint none of them holds, and the operator's intent is lost with no record.
- A longer interval divides the waste by whatever factor you choose and leaves the console permanently wrong. It hides the symptom that would otherwise have found the bug.
- Recording the rejection against the desired version terminates the loop, because the comparison is no longer value equality: it is "has this device acknowledged version N". A rejection is an acknowledgement, with a reason, and it belongs on an operator error queue.
What would have to be true for it to self-heal
Two things, and the design has neither. Every desired-state document needs a monotonic version, and every device needs a vocabulary for "I received version N and refused it, because the value is out of range for my model". Add a capability schema per model and firmware revision, validated before the write is accepted, and the bad bulk apply is rejected at the console for 40,000 devices in one response.
When this is over-engineering
A fleet of one hardware model with five attributes and a few hundred units does not need capability schemas or version negotiation. Value comparison plus a human who notices is adequate until either the model count or the device count grows, and the threshold in practice is the first time two firmware revisions accept different ranges for the same attribute.