No-Code SaaS Automation Platform  ·  View 21 of 21  ·  Assurance

Failure Classes and Their Handling

Eight classes, each with how it is detected, how the work is held, how it resolves, and what the author is told.

Editable source SVG draw.io All views
Detect Hold Resolve What the author sees Transient 5xx Timeout or reset Retry budget Same effect key Nothing Rate limit 429, Retry-After Park with release Resume, budget intact Running slowly Provider outage Sustained error rate Breaker, shed to floor Probe for recovery Provider is down Credential death 401 or revoked Park, then pause Author reconnects Reconnect your account Schema drift Schema validation Connector degraded Pin or update version One impact notice Ambiguous effect Acknowledgement lost Park if unsafe Read back, or ask An explicit choice Author logic error Publish-time check Fail fast Fix the binding A named cause Runaway automation 10x its baseline Throttle, then pause Author edits the loop Paused, with reason Assurance — Failure Classes and Their Handling Every row ends in a sentence the author can act on; a failure with no fourth column is an unfinished feature. v 1.0 · owner Integration Platform Architecture · date 2026-10

What the fourth column is for

  • A failure class with no author-facing sentence is an unfinished feature, not a support ticket. The fourth column is the acceptance test for every row.
  • Three rows resolve with the author doing nothing, and that is the point: transient errors, rate limits and provider outages are the platform's problem and should not reach a human who cannot act on them.
  • The two rows that end in a decision — credential death and the ambiguous effect — are the only ones where the platform genuinely cannot act alone, and both are escalated explicitly rather than guessed.

Assumptions

  • Rate limiting and credential death are continuous background events at 2.5 M workspaces, not incidents.
  • Author logic error is the single largest failure class by volume, which is why publish-time validation (view 04) is a gate rather than a warning.
  • Provider outages lasting hours occur several times a year per provider.

Not covered here

  • Worker loss, store loss and platform quota exhaustion are in the requirement's taxonomy but resolve structurally — through the ledger resume, synchronous replication and fair-share allocation — rather than through a detect-hold-resolve path, so they appear in views 07, 10 and 15 instead.