Customer 360 Enterprise Data Platform — Denodo on Azure  ·  View 39 of 39  ·  Assurance

Failure Modes

Every named way this breaks, what absorbs it, and the one class with no technical answer.

Editable source SVG draw.io All views
Source-side — expected, absorbed Source unavailable partial result, view 25 Source slow timeout, circuit opens API quota exhausted backoff, cache serves Schema drift base view breaks in CI Platform — degraded, not lost VDP pod loss retried, no session state Cache database loss cold reads, 3x latency MPP cluster loss falls back to the source Region loss RTO 30 min, RPO 5 min Correctness — the class that actually matters Stale cache after a merge reconciler bounds at 1 h Crosswalk corruption point in time, 7 days Wrong survivorship pick replayable from source ids Masking regression policy scan blocks release Programme — slow, and therefore easy to miss Uncertified view in use promotion gate blocks it Owner leaves, asset orphaned certification expires Extract copied to a spreadsheet usage telemetry only Denodo becomes the bottleneck the real programme risk same root cause concentration risk Failure Modes — What Breaks, What Absorbs It, and What Is Left Over Risk / gap synchronous failure / alternate The last box has no technical mitigation. Every consumer depending on one logical layer is the price of the design, and it is managed by workload isolation and honest capacity planning rather than by architecture. v 1.0 · owner Data & AI Global Practice · date 2026-09

Three classes, three responses

  • Source-side failures are expected and absorbed — partial results, circuit breakers, backoff, and CI catching schema drift before runtime does.
  • Platform failures degrade rather than lose. No component holds unique state except the crosswalk, and that one is replicated and point-in-time recoverable.
  • Correctness failures are the class that matters, and each has a named bound: an hour for stale caches, seven days for crosswalk recovery, a replay for wrong survivorship, a build failure for masking regressions.

The one with no technical mitigation

  • Denodo becoming the bottleneck for every consumer in the estate. That concentration is the price of a single logical layer, and it is managed by workload isolation, honest capacity planning and a licence model that is reviewed against actual concurrency — not designed away.
  • Second on that list: an extract copied into a spreadsheet. Usage telemetry narrows it and nothing closes it.

What is not on this page

  • Ordinary Azure platform failures — a zone loss, a managed service restart — which the deployment on view 28 already absorbs without a distinct response.