No-Code SaaS Automation Platform

Architecture Views

21 views, in reading order. Every view ships three ways: an HTML page, an SVG that re-opens in diagrams.net fully editable, and draw.io source.

A multi-tenant platform where a non-engineer connects their accounts in other people's SaaS products, describes "when this happens, do these things" by clicking, and the platform then runs it unattended for years. Read it in order: the boundary first, then who it is for, then the structure. The one decision everything else honours is that the step attempt — not the run — is the unit of durability, which is why the step ledger appears on almost every view from the fourth act onwards. Requirements: architecture/deliverables/use-cases/no-code-automation-platform/ask.md.

Context and scope

What sits inside the boundary, and what belongs to somebody else. Almost every hard requirement in this set exists because the interesting dependencies are not ours.

People and journeys

Who the platform is for, and what each of them gets to do. The second journey is the one that justifies the rest of the set.
03 The people who build and depend on automations Automation author Goal — I work in ops, sales or finance. Make the busywork between my tools disappear without asking engineering, and let me trust it is still running next quarter. Core journeys Build and publish one 12 M live Repair one that stopped Replay a run that failed Workspace admin Goal — I own IT and security here. Know which of our SaaS accounts this platform can reach, on whose authority, and cut any of it off in a minute. Core journeys Review connected accounts Revoke a connection Approve a shared-connection automation The people and machines that keep it running Platform operator Goal — I am the SRE on call. Tell me within a minute whether the backlog is our fault or a provider's, and shed the right load either way. Core journeys Triage a rising backlog Trip a provider breaker Connector developer partner or in-house Goal — Ship a connector for my product and have its replay safety stated honestly rather than guessed. Core journeys Publish a connector version Deprecate an old version Provider API quota owner Goal — Not be hammered. Have one caller I can identify, allowlist and rate-limit predictably. Core journeys Receive governed traffic Signal backoff and be obeyed Scheduler clock-driven Goal — Fire the scheduled and polled triggers on time without aligning a million connections on the minute. Core journeys Advance a connection cursor Who It Is For, and What They Get To Do Person or role Journey / task Application we own External / third party Security / platform Goals are in each actor's own voice; the two journeys with ids get their own map. v 1.0 · owner Integration Platform Architecture · date 2026-10 Actors and Their Journeys Five parties, their goals in their own words, and the journeys each of them gets. Two have their own map. HTML page SVG draw.io

Structure

The parts, their layers, and the interfaces — catalogued by replay safety rather than by vendor.

Data

Which two stores are authoritative, which can be dropped and rebuilt, and which one store's loss is unrecoverable.

Runtime

What actually happens when an automation fires, including the rate limit and the ambiguous outcome.

Operations

Deploy, release, watch and recover — including the signal that nothing is happening.

Assurance

Why it is safe to hand this platform a key to your other SaaS products, and what is assumed to fail.

Architecture One-Pager

The problem, the shape of the answer, the decisions that carry it, and what a prototype should prove.

The step attempt, not the run, is the unit of durability: a run is a resumable state machine whose entire state lives in an append-only step ledger outside the worker, and no step may be attempted without a stable effect key.

Almost everyone who works in an office has built one of these without calling it software. A form submission lands, a row appears in a spreadsheet, a message appears in a chat channel. Nobody wrote code and nobody deployed anything, and it has run unattended for two years — until the Tuesday it posts the same message three times, or the month it quietly stops because somebody in IT revoked a token and no human was told, or the Black Friday morning when 900 submissions trickle through over forty minutes because the spreadsheet API started answering 429. The product is trivial to describe and the architecture is not, because the platform owns almost none of the systems it depends on. Its triggers come from products that mostly cannot push. Its effects land in products whose APIs it cannot change, inside quotas it cannot negotiate, using credentials somebody else can revoke at any moment, on behalf of authors who cannot read a stack trace. Every interesting failure in this system belongs to somebody else — and every one of them still arrives as a complaint about us.

Separate what happened from what was done about it, and make the second one resumable. An authenticated provider delivery, or a polled read against a durable per-connection cursor, is committed to a partitioned trigger event log before the provider is acknowledged — so the accept path owes the caller durability and nothing else, and a total execution outage still loses no trigger. Admission deduplicates the event into at most one run per automation version, then schedules it under a per-workspace fair share across four queue classes so a bulk import cannot starve a two-step automation. A step runner leases the run and advances it one step at a time, writing an intent to the step ledger before each outbound call and an outcome after it; the ledger, not the worker's memory, is the run's state, so a lost worker resumes at the last committed step rather than repeating completed effects. Every outbound call leaves through a quota governor that owns the rate-limit decision, honours the provider's own backoff headers as authoritative, exchanges a short-lived single-connection token with a credential custody service in a separate account, and exits through a published NAT range customers can allowlist. A rate limit parks the run with a release time rather than failing it. An ambiguous outcome is resolved by the action's declared replay-safety class: retried under the same effect key when the provider honours one, read back first when a natural key exists, and escalated to the author when neither is true. Credential death, schema drift and runaway automations are detected, classified into an author-facing taxonomy, held rather than dropped, and surfaced in a repair inbox — because the default failure of this platform is silence, and silence has no natural end.

What it is, and what it is not

A resumable state machine whose state lives in an append-only ledgerA worker that holds a run in memory and retries it from the beginning
At-least-once execution with at-most-once visible effect, where the provider allows itA blanket exactly-once promise the providers cannot support
A declared replay-safety class per action, authored and versionedOne retry policy applied uniformly to forty thousand different operations
A rate limit as a normal run state with a release timeA rate limit as an error that consumes a retry budget
Credential custody in a separate account, reached by token exchangeA secret store the execution plane can read
Trigger ingestion that owns the subscription lifecycle and watches for silenceA webhook endpoint that assumes no news is good news
A failure taxonomy written in the author's vocabulary with a repair pathA provider error string forwarded to somebody who cannot read it
Automation of other people's SaaS products on an end user's behalfOrchestration of our own services, which is an adjacent use case

The decisions that are the architecture

01The step attempt is the unit of durability

A retry, a park for a rate limit, a pause for a dead credential and a resume after a lost worker all become the same mechanism — moving a run between states in the ledger — instead of four special cases, and none of them repeats a step that already succeeded.

ADR-01

02The trigger event log is our own system of record

Retention stops being a storage preference and becomes a replay commitment: the log is the only thing a replay reads, so thirty days is a promise about how far back a repair can reach, not a backup policy.

ADR-02

03Accept and execute fail independently

The push endpoint acknowledges after the durable commit and before any execution, which is how a 250 ms acknowledgement fits inside every common provider delivery timeout while the run itself may take minutes — and why an execution outage loses no trigger.

ADR-03

04Every action declares its replay-safety class

Idempotent, checkable or unsafe is a stored, versioned property of the action and a hard release gate, because every execution guarantee downstream is derived from it and an unclassified action would silently degrade the platform's correctness promise.

ADR-04

05A retry reuses the effect key; a replay mints a new attempt

The difference between 'try again' and 'do it again' is made structural rather than procedural, so the same code path cannot accidentally deliver one when the author asked for the other.

ADR-05

06An unsafe ambiguous outcome is never retried automatically

Where the provider offers no idempotency key and no read-back, the platform cannot make the step safe, so it parks and asks rather than silently choosing a duplicate invoice or a missing one on the author's behalf.

ADR-06

07No runner calls a provider directly

The quota decision, the credential exchange and the egress address are one boundary, which is what keeps long-lived credentials out of the execution plane and makes the platform's rate-limit behaviour a property rather than a convention.

ADR-07

08A rate limit is a park, not a failure

429 is the ordinary case at this scale, not an incident: the run moves to a waiting state with a release time, consumes no worker, and does not spend the step's retry budget on somebody else's capacity planning.

ADR-08

09Egress is a published, stable address range

Customers and providers allowlist it, which makes the range a product commitment that cannot change without notice — and makes one workspace's abusive automation everybody else's problem, which is why runaway detection is in the same decision.

ADR-09

10Push where available, poll everywhere else, and own the subscription

Roughly a fifth of the catalogue can push. Subscription renewal is a component rather than a setting, because a subscription that expired unnoticed is this platform's most common invisible failure and it presents as health.

ADR-10

11A cursor advances only after durable commit

Advancing first is the one-line bug that produces a permanent, silent gap. The cursor store is small, hot, and the only store here whose loss cannot be recovered from the two authoritative ones.

ADR-11

12Credential custody is a separate account

The threat is not that we lose our own data; it is that a compromise of our execution plane becomes a compromise of the customers' other SaaS products. An account boundary is the blast-radius boundary.

ADR-12

13Refresh is single-flighted; credential death is terminal

A thousand concurrent runs on one connection must not produce a thousand refreshes and a provider-side rotation race — a self-inflicted credential death — and a revoked token is never retried, only parked and reported.

ADR-13

14Fair scheduling, with runaway detection

A bulk import is slowed, never permitted to starve another workspace, and an automation triggering itself is throttled within a minute rather than consuming a quota and a provider's goodwill.

ADR-14

15Idle must be nearly free

Most of twelve million automations fire rarely. Near-zero idle cost is an architectural constraint that rules out per-automation reserved capacity, dedicated workers by default, and any design that holds a connection open per automation.

ADR-15

16Every failure carries an author-facing cause

A failure class with no sentence the author can act on is an unfinished feature, not a support ticket — which makes the failure taxonomy a deliverable of the architecture rather than a property of the UI.

ADR-16

Why this should still be right in ten years

The cloud, the queue, the container runtime and most of the eight thousand connectors will be replaced inside this platform's life. These are the properties that should outlast them.

Externalised run state is older than any of this technology

Putting a long-running process's state in a durable store rather than a worker's memory is the same decision a transaction log, a saga and a workflow engine all make. The specific store will change; the property that a worker may die at any instant without repeating a completed effect will not, because it follows from the fact that the effects are in somebody else's system and cannot be rolled back.

The effect key outlives the protocol

Idempotency keys are an HTTP convention today and were a message-deduplication identifier before that. What endures is the requirement that an operation carry a caller-supplied identity so a redelivery can be recognised as one. Any future transport will need it for the same reason, and the platform's derivation from (run, step, logical attempt) is transport-independent.

Replay safety is a property of the provider, not of our code

Classifying actions by whether a retry is safe will stay necessary for exactly as long as the platform integrates systems it does not control — which is permanently. The class may be discovered automatically one day rather than declared, but something will still have to carry the answer, and the architecture's dependence on it will not change.

Being a good citizen of a quota is a durable constraint

Providers will always meter access, and the penalty for overshoot will always be worse than the cost of pacing. A governance point that owns the outbound decision survives every change of rate-limit algorithm, because what it encodes is that the decision belongs in one place rather than in every caller.

Silence will always be the hard failure

In a system whose job is to act on somebody else's behalf, the costly failure is not an error but an absence — and no amount of better tooling makes an absence announce itself. A liveness expectation per connection and a loop that closes on recovery are the answer now and will be the answer on whatever this is rebuilt on.

The author will not become an engineer

The product's premise is that a non-engineer can build this. That constrains the architecture permanently: every failure must be reducible to a classified cause and a repair action, which rules out designs whose internal states cannot be explained in a sentence.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and follow it to the thing that depends on it.

QualityTargetHow it is metView
Trigger ingest availability ≥ 99.99% monthly Stateless push and poll tiers across three AZs, durable commit before acknowledgement, independent of the execution plane 13
Push acknowledgement latency p99 ≤ 250 ms Signature check and durable log commit only; no execution, binding or credential work on the accept path 12
Execution plane availability ≥ 99.95% monthly Leased runs resumable from the ledger, so a worker or AZ loss is a resume rather than an outage 15
Start latency, push-triggered p50 ≤ 800 ms, p95 ≤ 3 s, p99 ≤ 15 s Admission reads the log continuously and leases against a per-workspace fair share across four queue classes 02
Polled trigger detection Within one interval + 30 s at p95 Jittered per-connection scheduling at 1, 5 or 15 minute entitled intervals, with change-rate backoff for quiet connections 13
Throughput 12,000 step attempts/s steady, 48,000/s burst Horizontally scaled runners; the two synchronous ledger writes per step are the binding constraint 15
Duplicate visible effects ≤ 1 per 1,000,000 attempts (idempotent and checkable classes) Deterministic effect key per attempt, read-back before retry where no key is honoured 14
Silent run loss Zero tolerated Every accepted event reaches a terminal state within 7 days or appears in the repair inbox with a named cause 18
Outbound rate-limit rate ≤ 0.1% of calls receive a 429 Per-connection and per-connector quota leases at the governor, with provider backoff headers authoritative 17
Unclassified failure share ≤ 2% of author-facing failures Typed taxonomy with a named next action per class; above the threshold the taxonomy itself is the defect 21
Durability RPO 0 for accepted events and committed ledger entries Synchronous three-AZ replication on both authoritative stores; RPO 60 s for counters and projections 10
Recovery RTO 5 min in-region; 2 h to rebuild run history Run index is a projection of the ledger and is rebuilt without a maintenance window, degraded but available 09
Tenant isolation No workspace > 5% of a shared pool for > 60 s Fair scheduling with per-workspace concurrency leases and separated queue classes 02
Credential revocation Effective within 60 s for new runs Custody marks the connection dead; tokens already issued are scoped to one run and expire with its deadline 20
Cost ≤ $0.55 per 100,000 step attempts; ≤ $0.004 per idle automation per month Serverless runners with no per-automation reservation; ledger write held to ≤ 18% of per-step cost 17

Scope

In scope

  • The automation definition as immutable versioned configuration — trigger, ordered steps, reference-based data binding, branching, per-step error policy and the connector versions it was authored against — with publish-time validation and a 60-second rollback
  • Connector manifests as versioned declarative configuration, with input and output schemas, authentication scheme, rate-limit profile and a declared replay-safety class, plus a sandbox for the exceptional case of custom code
  • Trigger ingestion in four classes — provider push, polled, scheduled and inbound catch-hook or mail — with subscription lifecycle ownership, per-connection cursors, delivery authentication, dedup and poison quarantine
  • A durable partitioned trigger event log as the platform's own system of record for what happened, retained as the replay window
  • Run admission with per-event dedup, per-workspace fair scheduling, plan quotas and separated queue classes
  • Durable step execution against an append-only step ledger, with leases, effect keys, retry with backoff and jitter, run deadlines, bounded loops and checkpointed fan-out
  • Egress governance: per-connection and per-connector quota leases, provider backoff honoured as authoritative, park-not-fail on rate limits, per-provider circuit breaking and a published egress range
  • Credential custody as a separate plane: narrowest scopes, per-workspace encryption, single-flight refresh, short-lived single-connection tokens, credential-death detection and revocation
  • Author-facing visibility: classified run history, a repair inbox grouped by cause, author-initiated replay with the duplicate consequence stated, and notification on durable failure classes
  • Multi-tenancy, plan quotas, runaway detection and per-workspace data classification driving retention, residency and redaction

Explicitly out of scope

  • Generic DAG orchestration of the organisation's own internal services, which is the distributed-workflow-orchestration-platform use case in this practice
  • Outbound webhook fan-out to subscriber endpoints, which is the webhook-delivery-service use case
  • Request-level idempotency and deduplication as a standalone service, which is the idempotency-and-dedup-service use case
  • The connectors themselves beyond the manifest contract and the sandbox they run in — building eight thousand integrations is a programme, not an architecture
  • The authoring UI's interaction design, beyond the structural requirements it imposes (reference binding, publish validation, confirmed test writes)
  • Billing, pricing and plan definition beyond the quota and fairness mechanics the architecture must enforce
  • The providers' own systems, their uptime, their schemas and their rate-limit policies, all of which are constraints rather than components

What a four-week prototype should prove

The prototype's job is to falsify the three central claims — that step-level durability really does make a worker loss invisible in somebody else's system, that the effect key genuinely prevents a duplicate against a real provider, and that parking on a 429 is cheaper than failing — on three connectors and a few thousand runs, not to build a platform.

  1. Three real connectors chosen for contrast: one that pushes and honours an idempotency key (idempotent), one that must be polled and offers a read-back by natural key (checkable), and one that neither pushes nor keys — a chat or mail action (unsafe), so all three execution contracts are exercised rather than assumed
  2. A step ledger with intent-before-call and outcome-after-call, leases, and resume-at-last-committed-step, on one shared queue with a per-workspace concurrency lease
  3. A quota governor in front of all three connectors, honouring Retry-After, with a parked-release queue separate from the interactive one
  4. Credential custody as a separate process with per-connection single-flight refresh and run-scoped token issuance — the account boundary can wait, the interface cannot
  5. A trigger path per connector: signed push for the first, a durable per-connection cursor advanced only after commit for the second
  6. A classified failure taxonomy with a repair inbox, covering only the six classes the three connectors can actually produce
  7. Cost instrumented per 100,000 step attempts with the ledger writes broken out as their own line
  • Runners are killed mid-step at a measured rate under sustained load: the acceptance criterion is zero duplicated and zero skipped effects verified against the providers own records, not against our logs — this is the falsification test for ADR-01 and ADR-05 together
  • A deliberate 429 storm against the polled provider: parking must hold the run without occupying a worker, the retry budget must be untouched, and throughput must recover at the providers stated release time rather than at a multiple of it
  • The unsafe action is driven to a genuine ambiguity by severing the connection after the request is sent: the run must park with an accurate description of what may or may not have happened, and no automatic retry may occur
  • A credential is revoked at the provider mid-run: measure the elapsed time from revocation to the author being told, and whether the parked backlog is still worth releasing when it finally is
  • The cursor store is wiped for the polled connection: measure exactly how many duplicates or how large a gap results, because that number decides how durable the cursor store has to be and is the open half of Core Architecture Question 1
  • One workspace enqueues 50,000 runs while another runs a two-step automation: the second workspaces start latency must stay inside its budget, which is the only honest test of the fairness claim
  • A thousand concurrent runs hit an expired access token on one connection simultaneously: exactly one refresh must reach the provider, and the connection must survive — the falsification test for ADR-13

Open risks, carried rather than hidden

RiskIf it landsResponse
The step ledger's write rate becomes the platform's ceiling Two synchronously replicated writes per step at 12,000 steps/second is the hottest path in the system; if it cannot scale, the critical design decision has to be weakened to one write or to batched commits, which reintroduces the ambiguity it was adopted to remove Measure it first in the prototype with the ledger cost broken out; shard by run id; treat any proposal to batch the intent write as a change to ADR-01 rather than an optimisation
The replay-safety class is self-declared by connector developers whose incentive is to ship A wrong declaration is indistinguishable from a correct one until an effect is duplicated in a customer's system, and the platform's correctness promise silently degrades across the catalogue Make the class a hard release gate, test it in contract tests against provider sandboxes, and treat a confirmed misclassification as a connector incident with a deprecation, not a bug fix
Retaining thirty days of trigger payloads concentrates customer data from every connected product The platform becomes the custodian of a month of its customers' data across their whole SaaS estate — a breach, residency and retention surface far larger than its own records Core Architecture Question 5 is open: hash-and-reference with re-fetch is the alternative. Per-workspace no-payload-retention is offered now, with the repair path's limits stated honestly
The shared NAT egress range is a shared reputation One workspace's abusive or runaway automation can get the range rate-limited or blocked by a provider for every other workspace behind it Runaway detection within 60 s, per-connection and per-connector quota leases, and the ability to move a workspace to a separate egress range as an incident action
Partner connector code runs beside the execution plane The largest residual security risk in the set; a sandbox escape reaches the runs, and from there the token exchange Declarative connectors as the default and code as the exception; isolated function runtime with no credentials in scope, declared egress hosts and hard resource ceilings. Core Architecture Question 7 is whether to accept partner code at all
Parked backlogs become useless or dangerous while waiting for a human A week of parked runs released at once can flood a provider and perform work nobody wants any more, turning a repair into a second incident Rate-limited release, a run deadline that terminates rather than parking indefinitely, and Core Architecture Question 6 left open on whether release is automatic or confirmed
The quota governor is on the path of every outbound call A single hot dependency whose failure stops all effects platform-wide, and whose added hop is paid by every one of 12,000 steps a second Core Architecture Question 2 is open between central arbitration and per-connection leases; the lease design degrades to local decisions with bounded overshoot if the governor is unavailable

Architecture Decision Record

Why every component and every technology on these 21 views is what it is, and what each choice costs.

Sixteen decisions, grouped by the question they answer, each with the forcing question, the alternatives, what would flip the choice, and why it should outlast the technology it is realised on.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a multi-tenant automation platform with 2,500,000 active workspaces, 8,000 connectors exposing 40,000 actions and triggers, 12,000,000 enabled automations and roughly 1.8 billion step attempts a month; 4,000 trigger events a second accepted at steady state with a 4x ten-minute burst, and 12,000 step attempts a second rising to 48,000; 2,000,000 polled connections at an average five-minute interval, which is about 6,700 provider polls a second. Where the requirement supplied no number, one was invented and marked as an assumption in ask.md, section by section, so a reviewer can disagree with a figure and follow it to the decision that depends on it.

How to read a record

QuestionThe forcing question: why a decision was needed at all.
ContextThe requirement, the scale and the constraint that make it hard.
DecisionWhat this architecture does, stated so it can be checked.
How it is realised on AWSThe concrete mechanism: which service or package, configured how, in which subscription.
Options weighedChosen, rejected, deferred, or right elsewhere, with the reason for each.
ConsequencesWhat the choice buys and what it costs, both kept visible.
Choose differently whenThe conditions that would flip the decision for your system.
Why it holds up over timeWhat keeps the decision right as scale, staff and technology change.
LessonThe principle that transfers beyond this platform.

Decision map

The durability boundary 3

The decision that defines the architecture, and the two that follow from it directly.

ADR-01The step attempt, not the run, is the unit of durability ADR-02The trigger event log is the platform's own system of record for what happened ADR-03Acceptance and execution are separated by the log and fail independently

Effects in other people's systems 3

What a retry means when the thing being retried created an invoice somewhere we do not control.

ADR-04Every action declares a replay-safety class, and publication is gated on it ADR-05A retry reuses the effect key; a replay mints a new logical attempt ADR-06An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically

Living inside somebody else's quota 3

Rate limits we did not set, cannot negotiate, and are punished for ignoring.

ADR-07Every outbound call leaves through a quota governor; no runner calls a provider directly ADR-08A rate limit is a park with a release time, not a failure ADR-09Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation

Triggers from systems that cannot push 2

How the platform finds out that something happened, and how it notices that it has stopped finding out.

ADR-10Push where available, poll everywhere else, and own the subscription lifecycle as a component ADR-11A cursor advances only after durable commit, and absence of data is itself monitored

Credentials held on somebody's behalf 2

Custody of millions of delegated grants into other companies' products.

ADR-12Credential custody is a separate account; runners receive single-connection, single-run tokens ADR-13Refresh is single-flighted per connection, and credential death is terminal rather than retryable

Tenancy, fairness and the long tail 2

Twelve million automations, most of which do nothing most of the time.

ADR-14Fair scheduling on shared capacity, with runaway detection as part of the same decision ADR-15Near-zero idle cost is an architectural constraint, not an optimisation

Legibility 1

The author is not an engineer, and that is an architectural constraint rather than a UX preference.

ADR-16Every failure carries a classified, author-facing cause and a repair path

Technology by capability

Amazon Web Services was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases self-hosted open source carries 21 of 47 documented stacks and Microsoft Azure 13, with AWS on 9 — and Google Cloud, lower still at 7, took three of the seven most recent documents. Reaching for the same cloud each time teaches a service catalogue rather than architecture. The second is fit, in a weak sense that is worth stating honestly: this topic belongs to no cloud at all. Its hard parts — custody of delegated third-party credentials, a governor for quotas nobody hands you, a stable published egress identity, and a sandbox for partner code — are things every cloud leaves you to build. What AWS supplies well is the shape underneath: a very large fleet of short-lived, bursty, egress-heavy workers behind a fixed NAT range, with a partitioned durable log beside a high-write key-value store. Everything in ask.md is written vendor-neutrally — 'durable event log', not a product name — and the table below is where the neutral capability meets a specific service, with what would be used instead on another stack.

Open source This design
CapabilityChoiceOriginCredible alternativeWhy this oneRecord
Trigger event log — the platform's system of record for what happened Amazon MSK, partitioned by connection, 30 days retained as the replay window with tiered storage Amazon Web Services Confluent Cloud or self-managed Kafka; Azure Event Hubs with Capture; Google Pub/Sub with a replay subscription Per-connection ordering without global ordering, and a retention window that is a product promise about how far back a repair can reach ADR-02
Step ledger — run state, written twice per step DynamoDB, keyed (run_id, step_id, attempt), three-AZ synchronous replication, on-demand capacity Amazon Web Services Cloud Spanner or Bigtable; Azure Cosmos DB; CockroachDB or Cassandra self-hosted The hottest write path in the platform is a keyed append with no cross-row transaction; a relational store would buy consistency the access pattern does not need at a cost it cannot absorb ADR-01
Push ingestion endpoints ALB in front of Fargate services, per-connection signature verification, commit before acknowledgement Amazon Web Services Azure Front Door with Container Apps; Google Cloud Load Balancing with Cloud Run; Envoy on Kubernetes A stateless tier whose only obligation is to authenticate and durably commit inside the provider's delivery timeout ADR-03
Trigger cursors and subscription state DynamoDB, one small hot item per connection, advanced only after durable commit Amazon Web Services Azure Cosmos DB; Cloud Bigtable; Redis with AOF persistence, accepting the durability trade Two million connections of small, very hot, very frequently updated state — and the one store whose loss is not recoverable from the two authoritative ones ADR-11
Run admission and queue classes SQS, four queues by latency class, with per-workspace concurrency leases held in DynamoDB Amazon Web Services Azure Service Bus; Google Cloud Tasks; NATS JetStream or RabbitMQ self-hosted Separated queues make the fairness guarantee structural: a bulk import cannot share a queue with a two-step automation even by accident ADR-14
Step runners Fargate tasks, autoscaled on queue depth and age, holding no run state between steps Amazon Web Services Azure Container Apps; Cloud Run; Kubernetes with KEDA Serverless containers keep idle cost near zero across twelve million mostly-dormant automations, which rules out any per-automation reservation ADR-15
Quota governor and egress Fargate service owning per-connection and per-connector leases, behind NAT gateways with a published Elastic IP range Amazon Web Services Azure NAT Gateway with Container Apps; Cloud NAT with Cloud Run; Envoy with a rate-limit service on Kubernetes The quota decision, the token exchange and the egress identity belong at one boundary, and the address range is a customer-visible commitment ADR-07
Credential custody A separate AWS account running the custody service, with per-workspace KMS keys and envelope-encrypted material Amazon Web Services Azure Key Vault in a separate subscription; Cloud KMS in a separate project; HashiCorp Vault or OpenBao with a transit backend The blast-radius boundary is an account boundary, not a service boundary: a compromised task role in the platform account must not reach key material ADR-12
Connector sandbox AWS Lambda per invocation, no credentials in scope beyond the connection, declared egress hosts, hard CPU and memory ceilings Amazon Web Services Azure Functions; Cloud Functions; gVisor or Firecracker microVMs on Kubernetes; a WASM runtime for declarative-only connectors Partner code must not share a process, a filesystem or a network namespace with the run it serves ADR-04
Definition registry and tenancy Aurora PostgreSQL, immutable automation versions with a published-version pointer Amazon Web Services Azure Database for PostgreSQL; Cloud SQL or Spanner; PostgreSQL with Patroni self-hosted Publication is strongly consistent, low-volume and relational; the hot path never reads it without a cache ADR-05
Connector catalogue DynamoDB for manifests and version pointers, S3 for content-addressed artefacts Amazon Web Services Azure Cosmos DB with Blob Storage; Firestore with Cloud Storage; an OCI registry with MinIO Hot reads of varied-shape documents, with immutable artefacts addressed by digest so a pinned version is genuinely the same bytes ADR-04
Run history and analytics projections DynamoDB run index with a TTL by retention plan, S3 and Athena for cold history and cost analysis Amazon Web Services Azure Data Explorer; BigQuery; ClickHouse self-hosted Both are rebuildable from the ledger, so their schema can change without a migration and their loss is a degradation rather than an incident ADR-16

The decisions, and the alternatives that lost

The durability boundaryThe decision that defines the architecture, and the two that follow from it directly.

ADR-01

The step attempt, not the run, is the unit of durability

Accepted

When a worker dies halfway through a five-step automation that has already created an invoice and sent an email, what happens next — and who pays for it?

Context
A run in this platform is a sequence of effects in systems the platform does not own and cannot roll back. If run state lives in the worker's memory, the only recovery from a lost worker is to re-run from the beginning, which means re-performing every step that already succeeded. For a pipeline of pure computation that is a performance problem. Here it is a wrong answer in a customer's accounting system. The same argument applies to every other interruption: a rate limit that needs waiting out, a credential that died mid-run, a provider returning 503. Each of them needs the run to stop and later continue from exactly where it was, and none of them can be served by a design whose only verb is 'start over'. At 12,000 step attempts a second the cost of externalising that state is not theoretical either, which is why this is a decision rather than an obvious default.
Decision
A run is a resumable state machine whose complete state lives in an append-only step ledger outside the worker. Every step attempt writes its intent to the ledger before the outbound call and its outcome after, carrying the run, step, attempt number, resolved inputs, effect key, outcome and the provider's response identifier where one is returned. Run status is a projection of the ledger, never an independently updated field. A worker holds only a lease; losing it releases the run for another worker to resume at the last committed step. No step may be attempted without a stable effect key (ADR-05).
How it is realised on AWS
DynamoDB holds the ledger, keyed (run_id, step_id, attempt), replicated synchronously across three availability zones, with on-demand capacity and sharding by run id. Step runners are Fargate tasks that hold no state between steps; a run lease is a short-TTL item whose expiry makes the run eligible for another runner. The intent write and the outcome write are separate items rather than an update, so the ledger is genuinely append-only and the ambiguous case — intent present, outcome absent — is directly observable rather than inferred. Run status in the author-facing run index is materialised from the ledger asynchronously and can be rebuilt from it in full.
Options weighed
  • ChosenStep-level durability in an external append-only ledger: A park, a retry, a pause and a resume all become the same operation on the same structure, and no completed effect is ever repeated
  • RejectedRun-level durability: the worker holds the run and retries it from the start: Dramatically simpler and needs no ledger, but re-runs step one to retry step four — unacceptable when step one created an invoice
  • RejectedCheckpoint only at explicit author-declared boundaries: Pushes a correctness decision onto a non-engineer who has no way to reason about it, and leaves the default unsafe
  • RejectedA single mutable run row updated in place per step: Loses the intent-without-outcome signal that makes an ambiguous effect detectable, which is the case the whole design exists to handle
  • RejectedAn embedded workflow engine holding state in its own store: A reasonable alternative that moves the same decision behind a dependency; rejected because the effect-key semantics in ADR-05 are specific enough that owning the ledger is simpler than bending somebody else's
Consequences
What it buys
  • A worker loss, a rate-limit park, a credential pause and a transient retry are one mechanism — moving a run between states — rather than four special cases with four recovery paths
  • The ambiguous outcome is directly observable as an intent with no matching outcome, which is what makes ADR-06 implementable at all
  • Run history, the repair inbox and author-initiated replay are all reads of one structure, so they cannot disagree with each other
  • The run index can be dropped and rebuilt, which turns a schema change to author-facing history into a routine operation
What it costs
  • Two synchronously replicated writes per step attempt, which at 12,000 steps/second is the platform's hottest path and its primary scaling constraint
  • A priced component of every step: the ledger write is held to ≤ 18% of per-step cost, and that ceiling is a design constraint rather than a target
  • Latency added to every step by the intent write before the call, paid even on the overwhelming majority of steps that succeed first time
  • A lease mechanism with its own failure modes — a slow worker whose lease expires while its call is in flight is a duplicate risk that ADR-05 has to absorb
Choose differently when
Choose run-level retry instead when the effects are genuinely idempotent end to end, when runs are short enough that re-running from the start is cheap, or when the platform owns the systems the steps touch and can roll them back. A platform automating only its own internal services has all three properties, which is exactly why the adjacent distributed-workflow-orchestration-platform use case can make a different choice.
Why it holds up over time
Externalising long-running process state into a durable log is the same decision a transaction log, a saga and a workflow engine each make, and it has survived every generation of the technology underneath it. The store will be replaced; the property that a worker may die at any instant without repeating a completed effect follows from the effects being irreversible and somebody else's, which will not change.
LessonDecide what the unit of durability is before anything else, and derive the rest from it. If the thing you are retrying is irreversible and lives in a system you do not own, the unit cannot be the whole job.
Shown on views12 14 07 11
ADR-02

The trigger event log is the platform's own system of record for what happened

Accepted

When an author asks us to re-run last Tuesday's missed work, what do we read — our copy of the event, or the provider's current state?

Context
The providers are the system of record for what is true. They are not a reliable system of record for what happened: a webhook delivery is usually not retrievable afterwards, a polled change may have been superseded, and a deleted record simply is not there. If the platform keeps no copy, then a replay can only re-read the provider's current state, which means a repair performed a week after a credential died may run against data that has changed or vanished — or may be impossible because the credential is still dead, or because the provider is down. Keeping a copy makes replay exact and independent, and makes the platform the custodian of a month of its customers' data from every product they have connected.
Decision
Every accepted trigger event is durably committed to a partitioned, append-only event log before the provider is acknowledged, and retained for a declared replay window of 30 days. The log is the only thing a replay reads. Retention is therefore a recovery commitment — a promise about how far back a repair can reach — and not a storage preference.
How it is realised on AWS
Amazon MSK, partitioned by connection so per-connection ordering holds without requiring global ordering, with 30 days of retention using tiered storage so the window is affordable. Payload bodies above a size threshold are written to S3 and the log carries a reference, which keeps the log's partitions small and puts the residency and classification rules on object storage where they are easier to enforce. Admission reads the log continuously and deduplicates into at most one run per (automation version, event id).
Options weighed
  • ChosenRetain the payload for a declared replay window; replay reads the log: Replay is exact and independent of the provider's availability, credential state and current data
  • RejectedRetain only metadata and re-fetch the payload from the provider on replay: Holds almost nothing and materially reduces the breach surface, but makes replay silently run against changed data and impossible when the credential is dead — which is the most common reason a replay is needed
  • RejectedRetain nothing; a missed trigger is simply missed: Coherent and honest, and would remove a whole class of risk, but it deletes the repair journey the product is largely sold on
  • ChosenLet the workspace choose the window, including zero: Adopted as a per-workspace data-classification setting alongside the default, with the repair path's reduced capability stated plainly rather than discovered
Consequences
What it buys
  • A repair performed days later runs against exactly what arrived, not against what the provider happens to hold now
  • The accept path owes the provider durability and nothing else, which is what makes a 250 ms acknowledgement achievable
  • Admission, dedup and fair scheduling can all be rebuilt or re-run from the log without involving any provider
What it costs
  • Thirty days of customer data from every connected product concentrated in one platform — a residency, retention and breach surface far larger than the platform's own records
  • Retention is now a throughput obligation too: the log must be readable fast enough that a bulk replay does not starve live admission
  • A per-workspace no-retention setting creates a second, weaker repair path that has to be explained rather than hidden
Choose differently when
Choose metadata-plus-re-fetch when the payloads are highly sensitive, when providers offer a stable read-by-id that survives the retention period, and when replays are typically attempted within minutes rather than days. Choose to retain nothing when the product makes no repair promise at all.
Why it holds up over time
Separating 'what is true' from 'what happened' is a distinction every event-driven system eventually discovers, usually after trying to reconstruct history from current state and failing. The storage will change; the need for an independent record of the event, owned by the party that must act on it, will not.
LessonIf you promise to replay something, you must own a copy of it. A replay that reads the source is not a replay, it is a new run wearing the same name.
Shown on views09 10 13
ADR-03

Acceptance and execution are separated by the log and fail independently

Accepted

The execution plane is down. Does the platform tell the provider to go away, or take the event anyway?

Context
Provider push deliveries have short timeouts and limited, inconsistent retry behaviour; many providers retry two or three times and then drop the event permanently. A rejected delivery is therefore usually gone for good, and no amount of later recovery brings it back. Execution, by contrast, can be arbitrarily delayed without losing anything, because the work is still described in the log. The two sides of the platform have fundamentally different costs of failure, and giving them one availability target would either over-engineer execution or under-protect ingest.
Decision
The accept path authenticates the delivery, commits it durably to the event log, and acknowledges — and does nothing else. No binding, no credential resolution, no quota check and no execution happens before the acknowledgement. Ingest and execution carry separate availability targets (≥ 99.99% and ≥ 99.95%), scale independently, and a total execution outage loses no trigger.
How it is realised on AWS
ALB in front of stateless Fargate push services; the only work before acknowledgement is signature verification against the per-connection secret and the MSK produce with acknowledgement from all in-sync replicas. Admission runs as a separate consumer group entirely, so its failure, backlog or redeployment is invisible to the provider. Poll workers follow the same contract: read, commit, then advance the cursor (ADR-11).
Options weighed
  • ChosenAuthenticate, commit, acknowledge — nothing else on the accept path: The only design in which ingest availability is genuinely independent of everything downstream
  • RejectedExecute the first step synchronously before acknowledging: Gives the author a faster perceived start and ties the acknowledgement to credential and provider availability, which is precisely the coupling that loses events
  • RejectedValidate the payload against the automation's bindings before acknowledging: Catches author errors a few seconds earlier at the cost of making the accept path depend on the definition registry; the same error is caught at publish time instead (ADR-16)
  • RejectedReject deliveries when the execution backlog is beyond a threshold: Protects the platform by discarding exactly the events it exists to not lose; back-pressure is declared through widened latency budgets instead
Consequences
What it buys
  • A 250 ms p99 acknowledgement fits comfortably inside every common provider delivery timeout, while a run may legitimately take minutes or park for hours
  • Execution can be redeployed, scaled, or fail entirely without a single trigger being lost
  • Ingest is stateless apart from cursors, so it scales horizontally against burst without coordination
What it costs
  • The author's first feedback moves from the acknowledgement to the run history, so perceived latency has to be managed explicitly as a product surface
  • An automation whose definition is broken still accepts and logs events, producing runs that fail immediately and need a classified cause rather than a rejection
  • Two availability targets mean two operational postures, and the ingest tier must be genuinely over-provisioned rather than merely autoscaled
Choose differently when
Collapse the two when deliveries are reliably retried by the source for hours, or when the platform owns the source. A system reading from its own durable queue has no reason to separate acceptance from execution, because nothing is lost by refusing.
Why it holds up over time
The principle that the acceptance of work and the performance of work have different costs of failure, and therefore different availability targets, outlives any particular queue. It is the same reason a payment gateway acknowledges before settling.
LessonFind the step where failure is irreversible — usually the one where an external party gives up — and make it depend on as little as possible.
Shown on views02 12 13

Effects in other people's systemsWhat a retry means when the thing being retried created an invoice somewhere we do not control.

ADR-04

Every action declares a replay-safety class, and publication is gated on it

Accepted

Can this step be retried safely? And who is supposed to know the answer — the platform, the connector developer, or the author?

Context
The platform's execution guarantees cannot be stronger than what the providers support, and support varies wildly across 40,000 actions. Some accept a caller-supplied idempotency key. Some offer a read-back by natural key that reliably shows whether the effect landed. Many — notably 'post this message' and 'send this email', among the most-used actions in the product — offer neither, and after a timeout the platform genuinely cannot know what happened. A single platform-wide retry policy must therefore be wrong for most of the catalogue: safe enough for one class, and either duplicating or silently dropping effects for another.
Decision
Every action carries a declared replay-safety class — idempotent, checkable or unsafe — as a versioned field of the connector manifest. The class is a hard release gate: an action cannot be published without it. Execution semantics are derived from the class rather than configured separately, and the author-facing behaviour of the unsafe class is stated in the product.
How it is realised on AWS
The class lives in the connector manifest in DynamoDB alongside the schemas and rate-limit profile, versioned with the connector and pinned by each automation to a compatible range. The release pipeline rejects a manifest without it, and contract tests run against provider sandboxes rather than mocks, because the failure being guarded against is the provider's behaviour, which a mock cannot reveal. The step runner reads the class at bind time and selects the ambiguity-resolution path in ADR-06.
Options weighed
  • ChosenA declared, versioned per-action class, gated at release: Puts the answer where the knowledge is, and makes it reviewable, testable and pinnable
  • RejectedOne platform-wide retry policy: Simple and wrong for most of the catalogue: necessarily either duplicating or dropping, depending on which way it is tuned
  • RejectedInfer the class at runtime from provider responses: Attractive and unsound: the inference is only testable by causing the ambiguity it is meant to resolve, in a customer's system
  • RejectedLet the author choose per step: Asks a non-engineer a question they have no basis to answer, and the wrong answer is invisible until an invoice is duplicated
  • ChosenLet the author choose only the unsafe class's behaviour: Adopted in ADR-06: the author decides what to do about an ambiguity, never whether the ambiguity exists
Consequences
What it buys
  • The platform can state a different, honest guarantee per action instead of one promise that is false for most of them
  • The interface catalogue becomes organised by the property that actually matters, which makes the unsafe surface countable and reviewable
  • Pinning a connector version pins the guarantee, so a provider change cannot silently alter execution semantics for a published automation
What it costs
  • The class is self-declared by connector developers whose incentive is to ship, and a wrong declaration is indistinguishable from a right one until an effect is duplicated
  • A per-action field across 40,000 actions is a large surface to review and to keep honest as providers change
  • The product must explain a distinction users did not ask for, at the moment they are least interested in it
Choose differently when
Drop the classification when every integration target is under your control and can be made idempotent by fiat. An internal platform can mandate idempotency keys across its own services and needs none of this.
Why it holds up over time
For as long as a platform integrates systems it does not control, something must carry the answer to 'is a retry safe here'. The mechanism may one day be discovered rather than declared, but the dependency on the answer is permanent.
LessonWhen a guarantee depends on a third party's capability, make the capability an explicit, versioned, testable field — not an assumption spread through the code that uses it.
Shown on views08 14 16
ADR-05

A retry reuses the effect key; a replay mints a new logical attempt

Accepted

The author presses a button marked 'run it again'. Did they mean 'the first one may not have worked' or 'do it a second time on purpose'?

Context
Those two intentions produce opposite correct behaviours, and a platform that conflates them will sometimes duplicate an invoice and sometimes fail to send one. The distinction cannot live in a flag read at call time, because by then the two paths share the same code and the same provider call; it has to be built into how the key is derived, so that the difference is structural and cannot be got wrong by a caller.
Decision
Each step attempt carries a deterministic effect key derived from (run id, step id, logical attempt). A retry is a new physical attempt at the same logical attempt and therefore reuses the key, which is passed to the provider as an idempotency key wherever the action's class says it will be honoured. A replay is an author-initiated action that mints a new logical attempt and therefore a new key, is recorded in the ledger as a replay, and requires the duplicate consequence to be stated and confirmed first.
How it is realised on AWS
The key is a hash of the run id, step id and logical attempt number, computed by the runner before the intent write so that the ledger entry and the outbound call carry the same value. The governor forwards it in whatever header the connector manifest declares. The ledger's attempt key distinguishes physical attempts, while the logical attempt is a field on the entry — so 'how many times did we try' and 'how many times did the author ask for this' are separately countable.
Options weighed
  • ChosenKey derived from (run, step, logical attempt); replay increments the logical attempt: Makes the distinction structural: the same code cannot deliver a replay when a retry was meant
  • RejectedA random key per physical attempt: Defeats the entire purpose; every retry becomes a new effect at the provider
  • RejectedA key derived from the payload's content hash: Collides across two legitimately identical actions — two identical messages the author genuinely wanted twice — and silently drops the second
  • RejectedOne key per run: Cannot distinguish steps, so a provider honouring it would reject the second step of the same run
  • RejectedLet the connector supply its own key derivation: Flexible, and makes the platform's correctness depend on 8,000 independent implementations of the same subtle rule
Consequences
What it buys
  • A retry after a timeout, a lease expiry or a park cannot produce a second effect where the provider honours keys
  • 'Try again' and 'do it again' are visibly different in the ledger, so the audit answers which one happened
  • The key is transport-independent and survives a change of provider API or protocol
What it costs
  • The key must be minted before the intent write, so the derivation sits on the hottest path and cannot depend on anything slow
  • It is only useful where the provider honours it — the unsafe class carries the key and gets nothing for it, which has to be explained rather than assumed
  • A lease expiry during an in-flight call produces two physical attempts at the same logical attempt, which is safe only because the key is reused — making lease handling a correctness concern, not just a scheduling one
Choose differently when
A single key per operation is enough when the system performs one effect per job and never needs an intentional repeat. The two-level scheme earns its complexity only where a deliberate re-do is a product feature.
Why it holds up over time
Caller-supplied operation identity is an old idea with new names each decade — message dedup id, idempotency key, request id. What endures is that the caller, not the callee, has to assert what counts as the same operation.
LessonWhen two user intentions demand opposite behaviours, encode the difference in the data structure rather than in a parameter, so no code path can confuse them.
Shown on views11 12 14
ADR-06

An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically

Accepted

The request was sent, the connection dropped, and the provider offers no key and no read-back. Do we send it again or not?

Context
This is the one case where the platform genuinely cannot know what happened and cannot find out. Retrying risks a second email to a customer or a second message in a channel. Not retrying risks the thing never happening at all, silently. Both are wrong, and which is less wrong depends entirely on the action and the author's situation — a duplicate invoice is a serious problem, a duplicate internal reminder is not. The platform has no basis to choose, and choosing silently means being wrong without telling anybody.
Decision
For an action classified unsafe, an ambiguous outcome is never retried automatically. The run is parked in a state that names the ambiguity, the author is told exactly what may or may not have happened, and two explicit choices are offered: continue without retrying, accepting a possible gap, or retry and accept a possible duplicate. For checkable actions the platform reads back by natural key before retrying. For idempotent actions it simply retries under the same key.
How it is realised on AWS
The runner detects the ambiguity directly from the ledger's shape — an intent with no matching outcome — and branches on the manifest's class. The parked state carries a typed reason and the provider's last response if any, and surfaces in the repair inbox as its own category rather than mixed with ordinary failures. The run deadline still applies, so a parked-ambiguous run terminates into a named terminal state rather than waiting indefinitely for an author who will never look.
Options weighed
  • ChosenPark and escalate to the author with two explicit choices: The platform refuses to guess on the author's behalf about an effect in their own business
  • RejectedAlways retry: Produces duplicate invoices and duplicate customer emails, which is the failure customers escalate hardest and trust least
  • RejectedNever retry: Produces a silent gap, which is worse because nobody finds out — the exact failure mode journey 05 is about
  • RejectedLet the connector developer pick the default: Moves the guess one step away without improving the information available to make it
  • DeferredRefuse to publish automations whose first write is unsafe without acknowledgement: Attractive for high-stakes actions and left to Core Architecture Question 3; it raises a friction cost at authoring time that has not been measured
Consequences
What it buys
  • The platform never silently creates a duplicate effect in a customer's finance or communications system
  • The ambiguity is named and visible rather than resolved by a default nobody chose
  • The decision is made by the only party with the business context to make it
What it costs
  • A parked run waits for a human, and many authors will not look — the run deadline then terminates it, which is itself a form of the silent gap this decision tried to avoid
  • The product must explain a genuinely hard concept at a bad moment, and 'we do not know if your email sent' is a difficult sentence
  • Parked-ambiguous runs accumulate and need their own inventory, alerting and expiry policy
Choose differently when
Pick a default and apply it when the actions are low-stakes and homogeneous — an internal notifier can always retry, and nobody will mind. The escalation is worth its friction only where the effects are irreversible and matter commercially.
Why it holds up over time
Distributed systems will not stop producing ambiguous outcomes, and no protocol removes the case where the acknowledgement is lost and the operation cannot be queried. The lasting decision is whose judgement resolves it.
LessonWhen a system cannot know the answer and the consequences of guessing fall on someone else, surface the choice rather than picking a default and calling it a policy.
Shown on views14 21 12

Living inside somebody else's quotaRate limits we did not set, cannot negotiate, and are punished for ignoring.

ADR-07

Every outbound call leaves through a quota governor; no runner calls a provider directly

Accepted

Twelve thousand runners want to call the same provider at once. Who decides whether they may?

Context
Providers meter access per end-user account, per connector application credential, per region and per source address, and the penalties for overshoot escalate from a 429 to a suspended application credential that breaks every workspace at once. A decision made independently by each runner cannot respect a limit that is shared across runners, and the shared cases are the dangerous ones: a per-application limit is one scarce resource that every tenant is drawing on simultaneously. The same boundary is also the natural place to exchange a short-lived credential, because it is the last point before the call leaves the platform.
Decision
Every outbound third-party call passes a governance point that owns the quota decision; no step worker may call a provider on its own judgement. The governor enforces limits at each level the provider actually imposes them, honours provider backoff headers as authoritative over its own schedule, exchanges the short-lived single-connection token with custody, opens per-provider circuit breakers, and exits through the published egress range.
How it is realised on AWS
A Fargate service holding per-connection and per-connector leases in DynamoDB, in front of NAT gateways with a published Elastic IP range. Leases are checked out in blocks sized to the connection's observed rate so the common case is a local decision, with central arbitration reserved for per-application limits where tenants genuinely contend. Circuit breaker state is per provider and shared across the fleet, so one runner discovering an outage protects the rest.
Options weighed
  • ChosenA governance point owning the decision, with leases checked out in blocks: Correct for shared limits, and degrades to local decisions with bounded overshoot when the governor is slow
  • RejectedPer-worker local budgets with periodic reconciliation: Fastest and overshoots exactly when traffic is burstiest, which is when a provider is least forgiving
  • RejectedA central broker consulted synchronously on every call: Correct and puts a hard synchronous dependency plus a network hop on the hottest path in the platform
  • RejectedRely on the provider's 429 and back off reactively: Makes the provider's rate limiter our rate limiter, which works until the provider responds by suspending the application credential
Consequences
What it buys
  • A limit shared across tenants can actually be respected and fairly apportioned, which no local scheme achieves
  • Credential exchange, egress identity and the quota decision are one boundary, which is what keeps long-lived credentials out of the execution plane
  • Provider health becomes a first-class platform signal rather than something each runner rediscovers
What it costs
  • The governor is on the path of every outbound call and is the platform's hottest single dependency
  • Leases add a tuning problem: too small and they are a central broker with extra steps, too large and they overshoot during a burst
  • A cross-account token exchange per run adds latency, mitigated by caching for the run's duration at the cost of widening the revocation window
Choose differently when
Local budgets are sufficient when every limit is per end-user account, traffic is smooth, and the penalty for a modest overshoot is only a 429. Central arbitration earns its cost when tenants share one scarce credential.
Why it holds up over time
Providers will always meter, and the penalty for overshoot will always exceed the cost of pacing. Putting that decision in one place rather than in every caller survives any change of rate-limiting algorithm.
LessonA constraint that is shared cannot be enforced by parties acting independently. Put the decision where the constraint is, even when that means a hop on the hot path.
Shown on views07 12 19
ADR-08

A rate limit is a park with a release time, not a failure

Accepted

A provider says 'not now, try in ninety seconds'. Is that an error?

Context
At this scale a 429 is the ordinary case, not an incident — the platform makes billions of calls a month against quotas it did not set. Treating it as a failure has three bad consequences: it consumes the step's retry budget on something that is not a fault, it occupies a worker for the duration of a backoff, and it eventually exhausts the attempts and fails a run that was never actually broken. It also teaches the platform to treat the provider's own backoff signal as advice rather than instruction, which is how an application credential gets suspended.
Decision
A rate-limit response moves the run to a waiting state with a release time derived from the provider's own backoff header, which is authoritative in preference to the platform's schedule. A park consumes no worker and does not spend the step's retry budget. The connection's pacing adapts in response, so repeated parking feeds back into the lease size rather than simply repeating.
How it is realised on AWS
The governor returns a typed park outcome with a release timestamp; the runner writes it to the ledger and releases its lease. A dedicated parked-release queue class holds the run until its time, so parked work is isolated from interactive runs and cannot distort their queue-age signal. The platform's own SLO is that ≤ 0.1% of outbound calls receive a 429 at all — the governor's job is to make parking rare, not to make it cheap.
Options weighed
  • ChosenPark with a provider-derived release time, outside the retry budget: Treats the ordinary case as ordinary, and keeps the retry budget for actual faults
  • RejectedTreat 429 as a retryable error with exponential backoff: Spends the retry budget on capacity rather than faults, and holds a worker through every backoff
  • RejectedFail the run and let the author retry: Converts a routine, self-resolving condition into manual work for someone who cannot influence it
  • RejectedUse our own backoff schedule and ignore Retry-After: Substitutes our guess for the provider's statement about its own capacity, which is both rude and usually wrong
Consequences
What it buys
  • Throughput recovers at the provider's stated time rather than at an arbitrary multiple of it
  • The retry budget stays meaningful as a signal about faults, which keeps the failure taxonomy honest
  • Parked runs occupy no compute, so a provider-wide slowdown costs storage and patience rather than capacity
What it costs
  • A parked-release queue class with its own scheduling, inventory and alerting
  • Parked runs can accumulate into a backlog that fires in a burst at release, which needs its own rate-limited release path
  • The author experiences a slow automation with no error, which must be explained in the run history as 'running slowly' rather than left blank
Choose differently when
Treat rate limits as errors when they are genuinely exceptional and indicate misconfiguration rather than capacity — a well-provisioned internal API where a 429 means somebody has a bug.
Why it holds up over time
Backpressure from a dependency is information, not failure. Systems that learn to wait on it outlive systems that learn to retry through it, and that has been true of every protocol that has ever had a busy signal.
LessonDistinguish 'this is broken' from 'not right now'. Budgets, alerts and retries that conflate them will spend themselves on the wrong thing.
Shown on views12 21 17
ADR-09

Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation

Accepted

A customer's security team wants to allowlist us. What address do we give them, and what happens when a neighbour abuses it?

Context
Enterprise customers and many providers restrict inbound API access by source address, so a platform that cannot name its egress addresses cannot be adopted by them at all. Publishing a range makes it a commitment: it cannot change without notice, which constrains every future networking decision. It also makes the range a shared identity — providers rate-limit, throttle and block by source address, so a single runaway automation can degrade or sever access for every workspace behind it.
Decision
The platform egresses through NAT gateways with a published, stable address range that customers and providers may allowlist, and treats changes to it as a product change requiring notice. Because the range is shared, runaway detection, per-connection and per-connector quota leases, and the ability to move a workspace onto a separate egress range are part of the same decision rather than separate features.
How it is realised on AWS
NAT gateways with Elastic IPs per availability zone, published and versioned as documentation. Every outbound request carries a connector-identifying user agent and a correlation identifier a provider can quote in a support ticket. Residency-pinned workspaces egress from their own region's range. Runaway detection throttles an automation within 60 s of exceeding 10x its 7-day baseline, before a provider notices rather than after.
Options weighed
  • ChosenA published stable NAT range, with runaway controls as part of the same decision: Enterprise adoption requires it, and the shared-reputation risk it creates has to be mitigated in the same breath
  • RejectedEphemeral egress addresses: Operationally simplest and excludes every customer whose security policy requires an allowlist
  • RejectedA dedicated egress address per workspace: Removes the shared-reputation risk entirely and does not scale to 2.5 M workspaces at any sane cost
  • ChosenDedicated egress for workspaces that require or earn it: Adopted as both a plan feature and an incident response, on the same code path
Consequences
What it buys
  • Enterprise customers can allowlist the platform, which is a precondition for a large part of the market
  • A provider investigating traffic can identify the caller and quote a correlation id back, turning an opaque block into a conversation
  • Residency pinning extends naturally, because the egress identity is already per-region
What it costs
  • The range becomes immovable without a notice period, constraining future networking changes
  • One workspace's abuse is every workspace's problem, which raises the stakes on runaway detection considerably
  • Per-workspace egress as an escape hatch needs to exist before it is needed, not after
Choose differently when
Skip the published range when no customer requires allowlisting and providers do not restrict by address — a consumer-only product can egress from anywhere and avoid the shared-reputation problem entirely.
Why it holds up over time
Network identity as a trust signal has survived every change of protocol and will outlast the current generation of address-based controls, because the underlying need — naming who is calling — does not go away.
LessonA shared identity is a shared liability. If you publish one, you must also own the controls that stop one tenant spending the reputation of all of them.
Shown on views15 19 07

Triggers from systems that cannot pushHow the platform finds out that something happened, and how it notices that it has stopped finding out.

ADR-10

Push where available, poll everywhere else, and own the subscription lifecycle as a component

Accepted

Four fifths of the catalogue will not tell us when something changes. How does the platform find out, and how does it notice that it has stopped finding out?

Context
Provider push is cheap, fast and available for roughly a fifth of the catalogue, and it brings subscription lifecycle, signature verification, duplicate deliveries and — most damagingly — silent subscription death, where a subscription expires or is revoked and the platform simply stops receiving events with no error anywhere. Polling works everywhere, costs linearly in connections, and at 2,000,000 connections on an average five-minute interval amounts to roughly 6,700 provider polls a second that somebody has to pay for and that providers have to absorb.
Decision
Both, treated as architecturally different rather than hidden behind one abstraction. Push is preferred wherever the provider supports it, and the platform owns the subscription lifecycle as a component: create, renew before expiry, detect silent death, re-create. Polling adapts its interval per connection to the provider's quota profile, the plan's entitlement and the observed change rate, with declared jitter, and the initial poll establishes a watermark rather than emitting an account's entire history.
How it is realised on AWS
A subscription manager tracks every push subscription's renewal deadline and renews ahead of it; a liveness expectation per connection raises a quiet-period alarm when a source has emitted nothing for longer than its declared quiet period, which is wired off the event log so it watches for the absence of data rather than the health of the poller. Poll workers are Fargate tasks reading cursors from DynamoDB with per-connection jitter, backing off for quiet connections and batching provider queries where the API allows.
Options weighed
  • ChosenHybrid, with the subscription lifecycle owned explicitly: The only design that covers the catalogue while giving the fifth that can push the latency it deserves
  • RejectedPoll everything, for uniformity: One mechanism and one failure taxonomy, at the cost of latency on the connectors where latency is most visible and a much larger polling bill
  • RejectedPush only, and omit connectors that cannot push: Would cut the catalogue — the growth ceiling — by roughly four fifths
  • RejectedTreat polling as a fallback inside one trigger abstraction: Hides two genuinely different failure taxonomies behind one interface, which is how a dead subscription gets mistaken for a quiet account
Consequences
What it buys
  • Push connectors get seconds of latency and polled connectors get a declared, plan-appropriate interval, instead of both getting the worse of the two
  • A subscription that expires is detected and re-created rather than becoming an invisible outage
  • Jitter prevents the platform from aligning a million connections on the minute boundary against one provider
What it costs
  • Two freshness models, two dedup mechanisms and two failure taxonomies behind one product promise
  • Polling infrastructure is assumed at ≤ 25% of total platform compute cost, attributed to the connectors that lack push support
  • Subscription lifecycle is ongoing work that scales with the catalogue and with provider API churn
Choose differently when
Poll everything when the catalogue is small and latency expectations are loose; push only when every integration target is modern and under contract. The hybrid is the price of a long tail.
Why it holds up over time
The split between systems that notify and systems that must be asked has outlived several generations of integration technology and shows no sign of closing, because it reflects the providers' priorities rather than a technical limitation.
LessonDo not unify two mechanisms whose failure modes differ. One abstraction over both hides exactly the distinction operators need when something stops working.
Shown on views13 08 17
ADR-11

A cursor advances only after durable commit, and absence of data is itself monitored

Accepted

What is the smallest piece of state in this platform, and what happens when we lose it?

Context
A polled connection's cursor — a watermark, an updated-at bound or a page token — is a few bytes per connection. It is also the only state in the architecture whose loss cannot be recovered from the two authoritative stores: the event log knows what arrived, and the ledger knows what was done, but neither knows what the provider would have returned. Lose a cursor forward and there is a permanent silent gap; lose it backward and there is a flood of duplicates. Worse, the ordinary failure here produces no error at all — a stuck cursor and a healthy quiet account look identical.
Decision
A cursor advances only after the events it covers are durably committed to the event log, never before. The cursor and subscription store is treated as a distinct data zone with its own durability posture rather than grouped with other operational state. Every connection carries a liveness expectation, and an unexpectedly quiet source raises an alarm — the platform monitors the absence of data, not only the health of the components that fetch it.
How it is realised on AWS
DynamoDB, one small item per connection, written after the MSK produce is acknowledged by all in-sync replicas. The quiet-period expectation is derived from the connection's own observed history rather than a global constant, so a genuinely low-traffic connection does not alarm. The alarm is evaluated off the event log, so it fires whether the cause is a dead subscription, a stuck cursor, a paused automation or a provider that has quietly stopped returning results.
Options weighed
  • ChosenAdvance after commit, with a per-connection liveness expectation: Guarantees at-least-once at the cost of occasional duplicates, which ADR-05's effect key absorbs
  • RejectedAdvance before commit: A one-line difference that produces a permanent, silent, unrecoverable gap — the worst failure mode available to this platform
  • DeferredMake the cursor store as durable as the event log: Removes the residual risk and roughly doubles the cost of the hottest small store; left open as the second half of Core Architecture Question 1
  • RejectedReconstruct cursors from the event log after a loss: Recovers the last event seen but not the provider's pagination position, so it narrows the gap without closing it
Consequences
What it buys
  • A crash between read and commit costs a duplicate, which the effect key makes harmless, rather than a gap, which nothing makes harmless
  • A stuck cursor, a dead subscription and a silently broken provider all surface through one alarm that watches for silence
  • The riskiest store in the architecture is named as such, which is the precondition for anyone treating it carefully
What it costs
  • Duplicates are now routine on the poll path and the dedup window has to be sized for them
  • A per-connection liveness baseline is state about state, and it has to be learned rather than configured across two million connections
  • The residual risk is real and unresolved: a cursor store failure still produces duplicates or a gap, and the choice between them is not yet made
Choose differently when
Advancing before commit is defensible only when the source can be fully re-read cheaply and idempotently at any time, which makes the cursor an optimisation rather than state.
Why it holds up over time
Checkpoint-after-commit is the same rule as a consumer offset in a log, a replication position, or a file-read watermark. It will outlive this platform because it follows from the ordering of durability, not from any technology.
LessonFind the smallest piece of state whose loss cannot be reconstructed from anything else, and design for that one first. It is rarely the biggest store, and it is usually the one nobody is watching.
Shown on views13 10 17

Credentials held on somebody's behalfCustody of millions of delegated grants into other companies' products.

ADR-12

Credential custody is a separate account; runners receive single-connection, single-run tokens

Accepted

Our execution plane is compromised. How much of our customers' other SaaS estate goes with it?

Context
The platform holds delegated grants into millions of customers' CRMs, mailboxes, file stores and finance systems. Its own data is not the prize; the access is. The execution plane is simultaneously the largest attack surface in the system — it binds untrusted provider responses, runs partner connector code, and makes outbound calls to arbitrary hosts — and the component that needs credentials most often. Those two facts pull in opposite directions, and resolving them by convenience means a single compromised task role reads every refresh token the platform holds.
Decision
Credential custody runs in a separate cloud account with its own key material, reachable only through a narrow exchange interface. No long-lived credential is ever available to a step worker. A worker receives a short-lived token scoped to one connection and one run, exchanged at the egress boundary (ADR-07); the refresh token never leaves custody. Material is encrypted under a per-workspace key, and plaintext never appears in logs, run history, error messages or the ledger.
How it is realised on AWS
A separate AWS account running the custody service on Fargate, with per-workspace KMS keys and envelope-encrypted ciphertext. Cross-account access is by a narrowly scoped role permitting only the exchange operation, never a read. The connection's state and scopes stay in the platform account so the control plane can query them; only the ciphertext and key reference live in custody. Every exchange is written to an audit record retained seven years independently of the workspace's retention setting.
Options weighed
  • ChosenA separate account, with exchange-only access and run-scoped tokens: The blast-radius boundary matches the account boundary, which is the only boundary an attacker cannot talk their way across with a stolen role
  • RejectedA managed secret store in the same account, read by runners: Far simpler and makes a compromised runner equivalent to a compromise of every customer's connected products
  • RejectedA separate service in the same account: Better than nothing and still inside one IAM blast radius, which is the thing being defended against
  • RejectedHold no credentials: ask the author to re-authorise each run: Eliminates the risk and the product, which exists precisely to run unattended
Consequences
What it buys
  • A compromised execution plane yields run-scoped tokens for connections currently executing, not the estate
  • Per-workspace keys make crypto-deletion of one workspace's credentials a key operation rather than a data sweep
  • The audit of credential use is kept outside the workspace's own control, so it cannot be shortened by the party it holds accountable
What it costs
  • A cross-account hop on the path of every step that touches a provider, mitigated by per-run caching at the cost of a wider revocation window
  • Two accounts to operate, deploy and keep in step, with their own failure and permission-drift modes
  • Custody becomes a 99.99% dependency — every step needs it — which is a higher bar than the execution plane it serves
Choose differently when
A same-account secret store is proportionate when credentials reach only systems you already own, so a compromise gains the attacker nothing they did not already have.
Why it holds up over time
Separating custody of a credential from use of a credential is the oldest idea in access control and is independent of the technology implementing either. The boundary will move between accounts, enclaves or HSMs; the separation will not go away.
LessonSize the blast radius by what an attacker gains, not by what you store. If the value is access to other people's systems, the boundary has to be one your own compromised code cannot cross.
Shown on views19 20 07
ADR-13

Refresh is single-flighted per connection, and credential death is terminal rather than retryable

Accepted

A thousand runs on one connection hit an expired access token in the same second. What happens?

Context
Without coordination the answer is a thousand simultaneous refresh attempts against the provider. Many providers rotate the refresh token on use and invalidate the previous one, so concurrent refreshes race: one succeeds, the rest present a token that has just been invalidated, and the connection dies — a credential death entirely manufactured by the platform. The converse error is equally damaging: treating a genuine revocation as a transient fault and retrying produces a sustained authentication storm against a provider that has already said no, which is a reliable way to have an application credential suspended.
Decision
Access-token refresh is serialised per connection with single-flight semantics: one refresh proceeds, concurrent callers wait for its result. Credential death — revocation, expiry, password change, scope removal — is classified as non-retryable. Affected runs are parked, the automation is paused after a declared grace period, and the connection's owner is notified. The platform never retries an authentication failure as though it were transient.
How it is realised on AWS
Custody holds a per-connection lock for the duration of a refresh and returns the new token to all waiters. Failure classification distinguishes an expired access token (refresh and continue) from a dead grant (terminal), using the provider's error code as declared in the connector manifest rather than inferring from a status code alone. The terminal state feeds the automation health loop: classify, hold, notify, repair, observe.
Options weighed
  • ChosenSingle-flight refresh per connection; credential death terminal: Prevents the platform from causing the failure it is trying to recover from
  • RejectedLet each run refresh independently: Causes refresh-token rotation races and converts a routine expiry into a dead connection under load
  • ChosenRefresh proactively on a schedule: Adopted alongside, as an optimisation that reduces how often the single-flight path is contended; it does not remove the need for it
  • RejectedRetry authentication failures with backoff: Produces an authentication storm against a provider that has already declined, and risks suspension of the application credential
Consequences
What it buys
  • A connection under heavy concurrent load refreshes once, not a thousand times
  • A revoked credential fails fast and visibly instead of degrading slowly behind retries
  • Authentication traffic to providers stays proportional to connections rather than to runs
What it costs
  • A per-connection lock in custody, on the path of every step, with its own contention and timeout behaviour
  • Waiters block on a refresh they did not initiate, adding tail latency at exactly the moment a connection is busiest
  • Distinguishing 'expired' from 'revoked' depends on provider error semantics declared in the manifest, which is another field that can be wrong
Choose differently when
Independent refresh is harmless where providers do not rotate refresh tokens and tolerate concurrent refreshes. The single-flight requirement comes entirely from rotation semantics.
Why it holds up over time
Collapsing concurrent identical work into one operation is a pattern older than the web and applies wherever a shared resource is refreshed under load. Token rotation makes it a correctness requirement rather than an efficiency one, and rotation is becoming more common, not less.
LessonWhen a recovery action has side effects on shared state, concurrency control is part of correctness. An unsynchronised refresh is not a slow path, it is a bug.
Shown on views20 21 18

Tenancy, fairness and the long tailTwelve million automations, most of which do nothing most of the time.

ADR-14

Fair scheduling on shared capacity, with runaway detection as part of the same decision

Accepted

One workspace imports fifty thousand rows. What happens to everybody else's two-step automation?

Context
The workload is extremely uneven: most workspaces run a handful of automations occasionally, and a few run bulk operations that enqueue tens of thousands of runs in a burst. On one shared pool with a single queue, the bulk workspace wins simply by arriving, and every other tenant's latency collapses. Worse, some of those bursts are not legitimate work at all but automations triggering themselves — an automation that writes to the sheet it is watching — which can consume a workspace's quota and a provider's goodwill in minutes.
Decision
Shared capacity with fair scheduling across workspaces and per-workspace concurrency leases, so no workspace holds more than 5% of a shared pool for longer than 60 s while others have queued work. Queue classes are separated by expected latency — interactive, retry, parked-release and bulk — so a backfill never shares a queue with a short automation. Runaway detection throttles an automation within 60 s of exceeding 10x its own 7-day baseline, then pauses it with notice.
How it is realised on AWS
Four SQS queues by latency class, with per-workspace concurrency leases in DynamoDB enforced at admission rather than at the runner, so an over-quota workspace never occupies a worker it will not be allowed to use. Runaway baselines are per automation and learned from its own history, because a rate that is pathological for one automation is normal for another. Dedicated capacity for workspaces that require it runs on the same code path, as a lease pool rather than a separate fleet.
Options weighed
  • ChosenShared pool, fair scheduling, separated queue classes, runaway detection: Best utilisation with a bounded tail, and keeps idle cost near zero (ADR-15)
  • RejectedOne shared pool, first-come-first-served: Simplest and gives the worst tail exactly when it is most visible
  • RejectedPools per plan tier: Simple isolation that strands capacity and still does not protect tenants from each other within a tier
  • RejectedDedicated workers per workspace: Ends noisy neighbours and breaks the near-zero idle cost that makes twelve million automations affordable
  • ChosenDedicated capacity as a plan feature on the same code path: Adopted for the workspaces that require it, without creating a second execution path to maintain
Consequences
What it buys
  • A bulk import is slowed rather than permitted to starve, and the guarantee is structural because the queues are physically separate
  • A self-triggering automation is caught in a minute, before it spends a provider's goodwill or the workspace's quota
  • Utilisation stays high, which is what keeps per-step cost inside its target
What it costs
  • Fair scheduling is real machinery with its own tuning, and its failure mode is a latency regression that is hard to attribute
  • Per-automation baselines are state that must be learned across twelve million automations and will misfire on genuinely spiky but legitimate work
  • If plan-differentiated latency is ever sold, fairness and the commercial promise are in direct conflict and the architecture has to encode which wins
Choose differently when
First-come-first-served is adequate when the workload is homogeneous and no tenant can enqueue orders of magnitude more than another. Dedicated pools win when tenant isolation is a compliance requirement rather than a performance one.
Why it holds up over time
Fair queueing across tenants on shared infrastructure is as old as time-sharing and has survived every change in what the shared resource is. The specific scheduler will change; the need to stop one tenant monopolising a shared pool will not.
LessonMake isolation structural rather than behavioural. Separate queues enforce a guarantee that a shared queue with good intentions cannot.
Shown on views02 17 15
ADR-15

Near-zero idle cost is an architectural constraint, not an optimisation

Accepted

Twelve million automations exist. Most of them will not run today. What do they cost?

Context
The business is the long tail: the median workspace has a few automations that fire occasionally, and a large fraction of the twelve million are effectively dormant. If a dormant automation costs anything meaningful — a reserved worker, an open connection, a provisioned queue, a polling slot it does not need — the economics fail at a scale the product must reach to be viable. This is a genuine architectural constraint rather than a cost optimisation, because it rules out whole classes of design at the point where they are chosen.
Decision
An enabled but dormant automation holds no reserved compute, no open connection and no worker slot, at an assumed cost of ≤ $0.004 per automation per month. This rules out per-automation reserved capacity, dedicated workers by default, per-automation long-lived connections and any design where an automation's existence rather than its activity drives cost. Poll cost must be sub-linear in connection count, achieved through interval adaptation, change-rate backoff for quiet connections and batched provider queries.
How it is realised on AWS
Fargate tasks autoscaled on queue depth and age, with nothing allocated per automation. Automations exist as rows in the definition registry and subscriptions or cursors in a small key-value store — bytes, not capacity. Poll scheduling adapts per connection, so a connection that has not changed in a month is polled at its plan's floor rather than its ceiling.
Options weighed
  • ChosenServerless workers, no per-automation allocation, adaptive polling: The only shape where twelve million mostly-dormant automations are economically viable
  • RejectedA reserved worker or process per active automation: Simplest mental model and fails on cost by orders of magnitude at this scale
  • RejectedLong-running connections per automation: Would reduce latency and makes dormant automations expensive in exactly the way the economics cannot absorb
  • DeferredArchive dormant automations to cold storage: Worth doing if registry cost becomes material; it is not, because a dormant automation is already only bytes
Consequences
What it buys
  • The long tail is viable, which is the business rather than a technical nicety
  • Capacity tracks actual work, so a quiet weekend costs nothing and a Monday burst scales up
  • No per-automation resource means no per-automation leak, exhaustion or cleanup path
What it costs
  • Cold-start latency on scale-up, paid by whichever runs arrive first after a quiet period
  • A per-task premium over long-running instances, accepted deliberately as the price of the constraint
  • Adaptive polling adds state and complexity to the one tier that would otherwise be simple
Choose differently when
Per-tenant allocation is fine when tenants number in the thousands and each is commercially significant. The constraint comes from the tail's length, not from any dislike of reserved capacity.
Why it holds up over time
Making cost proportional to work rather than to configured entities is the economic principle under every generation of elastic infrastructure, and it will outlast the particular runtime that delivers it.
LessonWhen the business model is a long tail, the cost of an idle entity is an architectural constraint. Decide it before the design, because it eliminates options rather than tuning them.
Shown on views15 17 07

LegibilityThe author is not an engineer, and that is an architectural constraint rather than a UX preference.

ADR-16

Every failure carries a classified, author-facing cause and a repair path

Accepted

A non-engineer's automation failed. What do they see, and what can they do about it?

Context
The product's premise is that someone who cannot read a stack trace can build and own this. That premise turns failure presentation into an architectural concern rather than a UI one: if the platform's internal states cannot be reduced to a sentence an author can act on, then either the states are wrong or the product is undeliverable. The default alternative — forwarding the provider's error — is worse than useless, because provider errors are written for the developer who called the API, and often describe conditions the author has no way to influence.
Decision
Every failure is classified into a small, stable, author-facing taxonomy, each class carrying a named next action. An unclassified error is never presented as the primary message, and an unclassified share above 2% of author-facing failures is treated as a defect in the taxonomy. A repair inbox groups parked and failed runs by cause with bulk retry, replay, discard and reconnect actions. The classification exists in the execution path, not in the presentation layer.
How it is realised on AWS
The runner writes a typed cause onto the ledger entry at the moment of failure, where the provider response and the step context are both still available; the run index projects it, and the repair inbox groups on it. Provider-specific error codes are mapped to platform classes in the connector manifest, so the mapping is versioned with the connector that knows about it. Redaction of credential material and author-declared sensitive fields happens before anything reaches run history.
Options weighed
  • ChosenA typed taxonomy written at failure time, with a named action per class: The only approach where the classification has access to the context needed to be accurate
  • RejectedForward the provider's error message: Cheapest and describes conditions in a vocabulary the author cannot act on, often about systems they do not administer
  • RejectedClassify in the presentation layer from stored error text: Brittle string matching, far from the context, and silently wrong whenever a provider rewords a message
  • RejectedRoute every failure to human support: Does not scale past a tiny fraction of 1.8 billion monthly step attempts, and is slower than the author fixing it themselves
Consequences
What it buys
  • An author can usually repair their own automation without support, which is the only workable path at this volume
  • The unclassified rate becomes a measurable quality signal for the taxonomy itself, so gaps surface rather than accumulate
  • Grouping by cause makes bulk repair possible: one reconnection fixes hundreds of parked runs at once
What it costs
  • A mapping from provider errors to platform classes, per connector, which is ongoing work and another field that can be wrong
  • The taxonomy must stay small to be legible and complete enough to cover reality, and those pull against each other continuously
  • Redaction before storage means the raw provider response is not always available for operator debugging, which has its own cost
Choose differently when
Forwarding raw errors is correct when the users are developers who called the API themselves and have the context to interpret them. The taxonomy's cost is justified only by the non-engineer premise.
Why it holds up over time
The requirement that a system explain its failures in terms of the actions available to the person reading will outlast every interface technology, because it follows from who the user is rather than from how the message is rendered.
LessonIf your users cannot act on your errors, your error handling is unfinished — and that is a property of where classification happens in the architecture, not of how the message is styled.
Shown on views21 18 05

Every package used, in one table

Terms used in a specific sense in this package, where ordinary usage would be ambiguous.

PackageWhat it isWhat it does hereConsidered instead
Effect key A deterministic identifier derived from (run id, step id, logical attempt), passed to a provider as an idempotency key where the action's class says it will be honoured. The mechanism that makes a retry harmless: the same key means the same operation, so a redelivery is recognised rather than repeated. "Idempotency key", which in this package means the provider-side field the effect key is carried in, not the key itself.
Replay-safety class A declared, versioned property of a connector action: idempotent (the provider honours a caller key), checkable (the effect can be read back by a natural key), or unsafe (neither). The input from which every execution guarantee is derived, and a hard release gate on publishing an action. "Retry policy", which in this package means the backoff schedule and attempt ceiling, a separate and much weaker idea.
Park A normal run state with a release time, entered on a rate limit or an unresolvable ambiguity. It consumes no worker and does not spend the retry budget. The state that lets the platform wait on somebody else's capacity without treating waiting as failure. "Failed", which is terminal and consumes the budget; a parked run has neither property.
Retry A further physical attempt at the same logical attempt, reusing the effect key. Recovery from a fault, with no new effect intended. "Replay", which is the opposite intention.
Replay An author-initiated new logical attempt, with a new effect key, re-performing the effect deliberately from the retained trigger event. Recovery of missed work, with a new effect explicitly intended and confirmed. "Retry", which must never produce a second effect.
Step ledger The append-only store of step attempts — intent before each outbound call, outcome after — keyed by run, step and attempt. The platform's system of record for run state and the point a lost worker resumes from. Run status is a projection of it. "Run history", which is the rebuildable author-facing projection, not the authority.
Trigger event log The durable partitioned record of every accepted trigger event, retained for the replay window. The platform's system of record for what happened, and the only thing a replay reads. The providers, which remain the system of record for what is true.
Connection One workspace's authorised grant into one provider account, carrying scopes, state and a reference to credential material held in custody. The unit of credential custody, of provider rate limiting, and of cursor and subscription state. "Workspace", which is the unit of tenancy, quota, billing and audit.
Quiet period The interval after which a connection that has emitted nothing is considered unexpectedly silent, derived from its own observed history. Turns an absence of data into a signal, which is the only way a dead subscription or a stuck cursor becomes visible. A health check on the poller, which reports success while delivering nothing.
Unsafe action An action for which the provider offers neither a caller-supplied idempotency key nor a reliable read-back. The class for which the platform cannot deliver at-most-once visible effect, and says so rather than guessing. "Dangerous" or "unsupported" — an unsafe action is fully supported; what is unavailable is the automatic resolution of an ambiguous outcome.
Open svg/<view>.svg or drawio/<view>.drawio in draw.io Desktop or at app.diagrams.net to edit. The SVG carries the diagram inside it, so it is both the picture and the source. This folder is self-contained — copy it whole and every link still resolves.