No-Code SaaS Automation Platform

Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views

Almost everyone who works in an office has built one of these without calling it software. A form submission lands, a row appears in a spreadsheet, a message appears in a chat channel. Nobody wrote code and nobody deployed anything, and it has run unattended for two years — until the Tuesday it posts the same message three times, or the month it quietly stops because somebody in IT revoked a token and no human was told, or the Black Friday morning when 900 submissions trickle through over forty minutes because the spreadsheet API started answering 429. This package is the platform behind that — the Zapier / Make / IFTTT / Power Automate class: an assumed 2,500,000 active workspaces, 8,000 connectors exposing 40,000 actions and triggers, 12,000,000 enabled automations and roughly 1.8 billion step attempts a month; 4,000 trigger events a second accepted at steady state and 12,000 step attempts a second executed, both bursting fourfold; and 2,000,000 polled connections at an average five-minute interval, which is about 6,700 provider polls a second.

21 views 21 HTML views21 SVG21 draw.io 2 documents Updated 2026-10-10
Architecture views

21 views, each in three formats.

Open a view to read it in full. Every SVG carries its diagram source inside it, so it opens in diagrams.net fully editable with no import step; the draw.io files are the same diagrams as plain source.

  1. 01
    System Context

    Who uses the platform, which third-party products it reaches, and what two adjacent documents own instead.

  2. 02
    High-Level Architecture

    The whole path in one line: a provider event becomes a durable record, becomes a leased run, becomes a governed call back out.

  3. 03
    Actors and Their Journeys

    Five parties, their goals in their own words, and the journeys each of them gets. Two have their own map.

  4. 04
    Journey — Building the First Automation

    The happy path, and the one phase where the author is genuinely frightened of the product.

  5. 05
    Journey — The Automation That Quietly Stopped

    The journey that justifies the rest of the set: a revoked token, a week of silence, and a human finding out before we told them.

  6. 06
    Layered Architecture

    Six layers, with the providers drawn as the bottom one to keep the dependency direction honest.

  7. 07
    Container View

    The deployable units and their AWS services, with credential custody in its own account.

  8. 08
    Interface Catalogue

    Everything coming in, everything going out — and the outbound side grouped by whether a retry is safe.

  9. 09
    Data Flow

    A provider payload becomes an authoritative record twice, and everything after that is derived.

  10. 10
    Data Architecture

    Five zones by ownership and rebuildability, including the one whose loss cannot be recovered from the other four.

  11. 11
    Data Model

    Eight entities. The composite key on step_attempt is where exactly-once visible effect actually lives.

  12. 12
    Critical Flow — One Run With a Rate Limit

    The ordinary case, drawn with a 429 in the middle of it, because a 429 is the ordinary case.

  13. 13
    Trigger Ingestion

    Push and poll as two different mechanisms with two different failure modes, not one abstraction.

  14. 14
    Step Lifecycle by Replay-Safety Class

    The same five stages, three times, because the guarantee the platform can offer depends entirely on what the provider supports.

  15. 15
    Deployment

    One region across three availability zones, with the egress range customers can allowlist.

  16. 16
    Delivery and Connector Release

    Two things ship on this pipeline — our code and other people's connectors — and only one of them we wrote.

  17. 17
    Observability

    Six signal families by pipeline stage — including the row most platforms of this kind leave out.

  18. 18
    The Automation Health Loop

    Detect, classify, hold, notify, repair, observe — the loop that exists because the default failure here is silence.

  19. 19
    Security Zones

    Six zones by decreasing exposure, and every crossing labelled. The custody zone is a separate account on purpose.

  20. 20
    Identity and Access Flow

    Grant, store, exchange, use — and then the 401 that ends it, which is where this flow earns its place.

  21. 21
    Failure Classes and Their Handling

    Eight classes, each with how it is detected, how the work is held, how it resolves, and what the author is told.

Documents

The written architecture, on the page.

The view index above is the map; this is the argument. The one-pager and the decision record are part of the deliverable, so they are printed here in full — each also opens as its own page with a table of contents.

Document 1 of 2 · 18 min read

Architecture One-Pager

No-Code SaaS Automation Platform · Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views

The step attempt, not the run, is the unit of durability: a run is a resumable state machine whose entire state lives in an append-only step ledger outside the worker, and no step may be attempted without a stable effect key.

Almost everyone who works in an office has built one of these without calling it software. A form submission lands, a row appears in a spreadsheet, a message appears in a chat channel. Nobody wrote code and nobody deployed anything, and it has run unattended for two years — until the Tuesday it posts the same message three times, or the month it quietly stops because somebody in IT revoked a token and no human was told, or the Black Friday morning when 900 submissions trickle through over forty minutes because the spreadsheet API started answering 429. The product is trivial to describe and the architecture is not, because the platform owns almost none of the systems it depends on. Its triggers come from products that mostly cannot push. Its effects land in products whose APIs it cannot change, inside quotas it cannot negotiate, using credentials somebody else can revoke at any moment, on behalf of authors who cannot read a stack trace. Every interesting failure in this system belongs to somebody else — and every one of them still arrives as a complaint about us.

Separate what happened from what was done about it, and make the second one resumable. An authenticated provider delivery, or a polled read against a durable per-connection cursor, is committed to a partitioned trigger event log before the provider is acknowledged — so the accept path owes the caller durability and nothing else, and a total execution outage still loses no trigger. Admission deduplicates the event into at most one run per automation version, then schedules it under a per-workspace fair share across four queue classes so a bulk import cannot starve a two-step automation. A step runner leases the run and advances it one step at a time, writing an intent to the step ledger before each outbound call and an outcome after it; the ledger, not the worker's memory, is the run's state, so a lost worker resumes at the last committed step rather than repeating completed effects. Every outbound call leaves through a quota governor that owns the rate-limit decision, honours the provider's own backoff headers as authoritative, exchanges a short-lived single-connection token with a credential custody service in a separate account, and exits through a published NAT range customers can allowlist. A rate limit parks the run with a release time rather than failing it. An ambiguous outcome is resolved by the action's declared replay-safety class: retried under the same effect key when the provider honours one, read back first when a natural key exists, and escalated to the author when neither is true. Credential death, schema drift and runaway automations are detected, classified into an author-facing taxonomy, held rather than dropped, and surfaced in a repair inbox — because the default failure of this platform is silence, and silence has no natural end.

What it is, and what it is not

  • A resumable state machine whose state lives in an append-only ledger — not A worker that holds a run in memory and retries it from the beginning
  • At-least-once execution with at-most-once visible effect, where the provider allows it — not A blanket exactly-once promise the providers cannot support
  • A declared replay-safety class per action, authored and versioned — not One retry policy applied uniformly to forty thousand different operations
  • A rate limit as a normal run state with a release time — not A rate limit as an error that consumes a retry budget
  • Credential custody in a separate account, reached by token exchange — not A secret store the execution plane can read
  • Trigger ingestion that owns the subscription lifecycle and watches for silence — not A webhook endpoint that assumes no news is good news
  • A failure taxonomy written in the author's vocabulary with a repair path — not A provider error string forwarded to somebody who cannot read it
  • Automation of other people's SaaS products on an end user's behalf — not Orchestration of our own services, which is an adjacent use case

The decisions that are the architecture

  1. The step attempt is the unit of durability (ADR-01) — A retry, a park for a rate limit, a pause for a dead credential and a resume after a lost worker all become the same mechanism — moving a run between states in the ledger — instead of four special cases, and none of them repeats a step that already succeeded.
  2. The trigger event log is our own system of record (ADR-02) — Retention stops being a storage preference and becomes a replay commitment: the log is the only thing a replay reads, so thirty days is a promise about how far back a repair can reach, not a backup policy.
  3. Accept and execute fail independently (ADR-03) — The push endpoint acknowledges after the durable commit and before any execution, which is how a 250 ms acknowledgement fits inside every common provider delivery timeout while the run itself may take minutes — and why an execution outage loses no trigger.
  4. Every action declares its replay-safety class (ADR-04) — Idempotent, checkable or unsafe is a stored, versioned property of the action and a hard release gate, because every execution guarantee downstream is derived from it and an unclassified action would silently degrade the platform's correctness promise.
  5. A retry reuses the effect key; a replay mints a new attempt (ADR-05) — The difference between 'try again' and 'do it again' is made structural rather than procedural, so the same code path cannot accidentally deliver one when the author asked for the other.
  6. An unsafe ambiguous outcome is never retried automatically (ADR-06) — Where the provider offers no idempotency key and no read-back, the platform cannot make the step safe, so it parks and asks rather than silently choosing a duplicate invoice or a missing one on the author's behalf.
  7. No runner calls a provider directly (ADR-07) — The quota decision, the credential exchange and the egress address are one boundary, which is what keeps long-lived credentials out of the execution plane and makes the platform's rate-limit behaviour a property rather than a convention.
  8. A rate limit is a park, not a failure (ADR-08) — 429 is the ordinary case at this scale, not an incident: the run moves to a waiting state with a release time, consumes no worker, and does not spend the step's retry budget on somebody else's capacity planning.
  9. Egress is a published, stable address range (ADR-09) — Customers and providers allowlist it, which makes the range a product commitment that cannot change without notice — and makes one workspace's abusive automation everybody else's problem, which is why runaway detection is in the same decision.
  10. Push where available, poll everywhere else, and own the subscription (ADR-10) — Roughly a fifth of the catalogue can push. Subscription renewal is a component rather than a setting, because a subscription that expired unnoticed is this platform's most common invisible failure and it presents as health.
  11. A cursor advances only after durable commit (ADR-11) — Advancing first is the one-line bug that produces a permanent, silent gap. The cursor store is small, hot, and the only store here whose loss cannot be recovered from the two authoritative ones.
  12. Credential custody is a separate account (ADR-12) — The threat is not that we lose our own data; it is that a compromise of our execution plane becomes a compromise of the customers' other SaaS products. An account boundary is the blast-radius boundary.
  13. Refresh is single-flighted; credential death is terminal (ADR-13) — A thousand concurrent runs on one connection must not produce a thousand refreshes and a provider-side rotation race — a self-inflicted credential death — and a revoked token is never retried, only parked and reported.
  14. Fair scheduling, with runaway detection (ADR-14) — A bulk import is slowed, never permitted to starve another workspace, and an automation triggering itself is throttled within a minute rather than consuming a quota and a provider's goodwill.
  15. Idle must be nearly free (ADR-15) — Most of twelve million automations fire rarely. Near-zero idle cost is an architectural constraint that rules out per-automation reserved capacity, dedicated workers by default, and any design that holds a connection open per automation.
  16. Every failure carries an author-facing cause (ADR-16) — A failure class with no sentence the author can act on is an unfinished feature, not a support ticket — which makes the failure taxonomy a deliverable of the architecture rather than a property of the UI.

Why this should still be right in ten years

The cloud, the queue, the container runtime and most of the eight thousand connectors will be replaced inside this platform's life. These are the properties that should outlast them.

  • Externalised run state is older than any of this technology. Putting a long-running process's state in a durable store rather than a worker's memory is the same decision a transaction log, a saga and a workflow engine all make. The specific store will change; the property that a worker may die at any instant without repeating a completed effect will not, because it follows from the fact that the effects are in somebody else's system and cannot be rolled back.
  • The effect key outlives the protocol. Idempotency keys are an HTTP convention today and were a message-deduplication identifier before that. What endures is the requirement that an operation carry a caller-supplied identity so a redelivery can be recognised as one. Any future transport will need it for the same reason, and the platform's derivation from (run, step, logical attempt) is transport-independent.
  • Replay safety is a property of the provider, not of our code. Classifying actions by whether a retry is safe will stay necessary for exactly as long as the platform integrates systems it does not control — which is permanently. The class may be discovered automatically one day rather than declared, but something will still have to carry the answer, and the architecture's dependence on it will not change.
  • Being a good citizen of a quota is a durable constraint. Providers will always meter access, and the penalty for overshoot will always be worse than the cost of pacing. A governance point that owns the outbound decision survives every change of rate-limit algorithm, because what it encodes is that the decision belongs in one place rather than in every caller.
  • Silence will always be the hard failure. In a system whose job is to act on somebody else's behalf, the costly failure is not an error but an absence — and no amount of better tooling makes an absence announce itself. A liveness expectation per connection and a loop that closes on recovery are the answer now and will be the answer on whatever this is rebuilt on.
  • The author will not become an engineer. The product's premise is that a non-engineer can build this. That constrains the architecture permanently: every failure must be reducible to a classified cause and a repair action, which rules out designs whose internal states cannot be explained in a sentence.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and follow it to the thing that depends on it.

Quality Target How it is met View
Trigger ingest availability ≥ 99.99% monthly Stateless push and poll tiers across three AZs, durable commit before acknowledgement, independent of the execution plane 13
Push acknowledgement latency p99 ≤ 250 ms Signature check and durable log commit only; no execution, binding or credential work on the accept path 12
Execution plane availability ≥ 99.95% monthly Leased runs resumable from the ledger, so a worker or AZ loss is a resume rather than an outage 15
Start latency, push-triggered p50 ≤ 800 ms, p95 ≤ 3 s, p99 ≤ 15 s Admission reads the log continuously and leases against a per-workspace fair share across four queue classes 02
Polled trigger detection Within one interval + 30 s at p95 Jittered per-connection scheduling at 1, 5 or 15 minute entitled intervals, with change-rate backoff for quiet connections 13
Throughput 12,000 step attempts/s steady, 48,000/s burst Horizontally scaled runners; the two synchronous ledger writes per step are the binding constraint 15
Duplicate visible effects ≤ 1 per 1,000,000 attempts (idempotent and checkable classes) Deterministic effect key per attempt, read-back before retry where no key is honoured 14
Silent run loss Zero tolerated Every accepted event reaches a terminal state within 7 days or appears in the repair inbox with a named cause 18
Outbound rate-limit rate ≤ 0.1% of calls receive a 429 Per-connection and per-connector quota leases at the governor, with provider backoff headers authoritative 17
Unclassified failure share ≤ 2% of author-facing failures Typed taxonomy with a named next action per class; above the threshold the taxonomy itself is the defect 21
Durability RPO 0 for accepted events and committed ledger entries Synchronous three-AZ replication on both authoritative stores; RPO 60 s for counters and projections 10
Recovery RTO 5 min in-region; 2 h to rebuild run history Run index is a projection of the ledger and is rebuilt without a maintenance window, degraded but available 09
Tenant isolation No workspace > 5% of a shared pool for > 60 s Fair scheduling with per-workspace concurrency leases and separated queue classes 02
Credential revocation Effective within 60 s for new runs Custody marks the connection dead; tokens already issued are scoped to one run and expire with its deadline 20
Cost ≤ $0.55 per 100,000 step attempts; ≤ $0.004 per idle automation per month Serverless runners with no per-automation reservation; ledger write held to ≤ 18% of per-step cost 17

Scope

In scope

  • The automation definition as immutable versioned configuration — trigger, ordered steps, reference-based data binding, branching, per-step error policy and the connector versions it was authored against — with publish-time validation and a 60-second rollback
  • Connector manifests as versioned declarative configuration, with input and output schemas, authentication scheme, rate-limit profile and a declared replay-safety class, plus a sandbox for the exceptional case of custom code
  • Trigger ingestion in four classes — provider push, polled, scheduled and inbound catch-hook or mail — with subscription lifecycle ownership, per-connection cursors, delivery authentication, dedup and poison quarantine
  • A durable partitioned trigger event log as the platform's own system of record for what happened, retained as the replay window
  • Run admission with per-event dedup, per-workspace fair scheduling, plan quotas and separated queue classes
  • Durable step execution against an append-only step ledger, with leases, effect keys, retry with backoff and jitter, run deadlines, bounded loops and checkpointed fan-out
  • Egress governance: per-connection and per-connector quota leases, provider backoff honoured as authoritative, park-not-fail on rate limits, per-provider circuit breaking and a published egress range
  • Credential custody as a separate plane: narrowest scopes, per-workspace encryption, single-flight refresh, short-lived single-connection tokens, credential-death detection and revocation
  • Author-facing visibility: classified run history, a repair inbox grouped by cause, author-initiated replay with the duplicate consequence stated, and notification on durable failure classes
  • Multi-tenancy, plan quotas, runaway detection and per-workspace data classification driving retention, residency and redaction

Explicitly out of scope

  • Generic DAG orchestration of the organisation's own internal services, which is the distributed-workflow-orchestration-platform use case in this practice
  • Outbound webhook fan-out to subscriber endpoints, which is the webhook-delivery-service use case
  • Request-level idempotency and deduplication as a standalone service, which is the idempotency-and-dedup-service use case
  • The connectors themselves beyond the manifest contract and the sandbox they run in — building eight thousand integrations is a programme, not an architecture
  • The authoring UI's interaction design, beyond the structural requirements it imposes (reference binding, publish validation, confirmed test writes)
  • Billing, pricing and plan definition beyond the quota and fairness mechanics the architecture must enforce
  • The providers' own systems, their uptime, their schemas and their rate-limit policies, all of which are constraints rather than components

What a four-week prototype should prove

The prototype's job is to falsify the three central claims — that step-level durability really does make a worker loss invisible in somebody else's system, that the effect key genuinely prevents a duplicate against a real provider, and that parking on a 429 is cheaper than failing — on three connectors and a few thousand runs, not to build a platform.

  1. Three real connectors chosen for contrast: one that pushes and honours an idempotency key (idempotent), one that must be polled and offers a read-back by natural key (checkable), and one that neither pushes nor keys — a chat or mail action (unsafe), so all three execution contracts are exercised rather than assumed
  2. A step ledger with intent-before-call and outcome-after-call, leases, and resume-at-last-committed-step, on one shared queue with a per-workspace concurrency lease
  3. A quota governor in front of all three connectors, honouring Retry-After, with a parked-release queue separate from the interactive one
  4. Credential custody as a separate process with per-connection single-flight refresh and run-scoped token issuance — the account boundary can wait, the interface cannot
  5. A trigger path per connector: signed push for the first, a durable per-connection cursor advanced only after commit for the second
  6. A classified failure taxonomy with a repair inbox, covering only the six classes the three connectors can actually produce
  7. Cost instrumented per 100,000 step attempts with the ledger writes broken out as their own line
  • Runners are killed mid-step at a measured rate under sustained load: the acceptance criterion is zero duplicated and zero skipped effects verified against the providers own records, not against our logs — this is the falsification test for ADR-01 and ADR-05 together
  • A deliberate 429 storm against the polled provider: parking must hold the run without occupying a worker, the retry budget must be untouched, and throughput must recover at the providers stated release time rather than at a multiple of it
  • The unsafe action is driven to a genuine ambiguity by severing the connection after the request is sent: the run must park with an accurate description of what may or may not have happened, and no automatic retry may occur
  • A credential is revoked at the provider mid-run: measure the elapsed time from revocation to the author being told, and whether the parked backlog is still worth releasing when it finally is
  • The cursor store is wiped for the polled connection: measure exactly how many duplicates or how large a gap results, because that number decides how durable the cursor store has to be and is the open half of Core Architecture Question 1
  • One workspace enqueues 50,000 runs while another runs a two-step automation: the second workspaces start latency must stay inside its budget, which is the only honest test of the fairness claim
  • A thousand concurrent runs hit an expired access token on one connection simultaneously: exactly one refresh must reach the provider, and the connection must survive — the falsification test for ADR-13

Open risks, carried rather than hidden

Risk If it lands Response
The step ledger's write rate becomes the platform's ceiling Two synchronously replicated writes per step at 12,000 steps/second is the hottest path in the system; if it cannot scale, the critical design decision has to be weakened to one write or to batched commits, which reintroduces the ambiguity it was adopted to remove Measure it first in the prototype with the ledger cost broken out; shard by run id; treat any proposal to batch the intent write as a change to ADR-01 rather than an optimisation
The replay-safety class is self-declared by connector developers whose incentive is to ship A wrong declaration is indistinguishable from a correct one until an effect is duplicated in a customer's system, and the platform's correctness promise silently degrades across the catalogue Make the class a hard release gate, test it in contract tests against provider sandboxes, and treat a confirmed misclassification as a connector incident with a deprecation, not a bug fix
Retaining thirty days of trigger payloads concentrates customer data from every connected product The platform becomes the custodian of a month of its customers' data across their whole SaaS estate — a breach, residency and retention surface far larger than its own records Core Architecture Question 5 is open: hash-and-reference with re-fetch is the alternative. Per-workspace no-payload-retention is offered now, with the repair path's limits stated honestly
The shared NAT egress range is a shared reputation One workspace's abusive or runaway automation can get the range rate-limited or blocked by a provider for every other workspace behind it Runaway detection within 60 s, per-connection and per-connector quota leases, and the ability to move a workspace to a separate egress range as an incident action
Partner connector code runs beside the execution plane The largest residual security risk in the set; a sandbox escape reaches the runs, and from there the token exchange Declarative connectors as the default and code as the exception; isolated function runtime with no credentials in scope, declared egress hosts and hard resource ceilings. Core Architecture Question 7 is whether to accept partner code at all
Parked backlogs become useless or dangerous while waiting for a human A week of parked runs released at once can flood a provider and perform work nobody wants any more, turning a repair into a second incident Rate-limited release, a run deadline that terminates rather than parking indefinitely, and Core Architecture Question 6 left open on whether release is automatic or confirmed
The quota governor is on the path of every outbound call A single hot dependency whose failure stops all effects platform-wide, and whose added hop is paid by every one of 12,000 steps a second Core Architecture Question 2 is open between central arbitration and per-connection leases; the lease design degrades to local decisions with bounded overshoot if the governor is unavailable

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 7 areas, each with the alternatives that lost and what the choice costs.

Document 2 of 2 · 61 min read

Architecture Decision Record

No-Code SaaS Automation Platform · Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views

The argument these decisions serve is summarised in the Architecture One-Pager.

Sixteen decisions, grouped by the question they answer, each with the forcing question, the alternatives, what would flip the choice, and why it should outlast the technology it is realised on.

Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a multi-tenant automation platform with 2,500,000 active workspaces, 8,000 connectors exposing 40,000 actions and triggers, 12,000,000 enabled automations and roughly 1.8 billion step attempts a month; 4,000 trigger events a second accepted at steady state with a 4x ten-minute burst, and 12,000 step attempts a second rising to 48,000; 2,000,000 polled connections at an average five-minute interval, which is about 6,700 provider polls a second. Where the requirement supplied no number, one was invented and marked as an assumption in ask.md, section by section, so a reviewer can disagree with a figure and follow it to the decision that depends on it.

How to read a record

  • Question: The forcing question: why a decision was needed at all.
  • Context: The requirement, the scale and the constraint that make it hard.
  • Decision: What this architecture does, stated so it can be checked.
  • How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
  • Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
  • Consequences: What the choice buys and what it costs, both kept visible.
  • Choose differently when: The conditions that would flip the decision for your system.
  • Why it holds up over time: What keeps the decision right as scale, staff and technology change.
  • Lesson: The principle that transfers beyond this platform.

Decision map

The durability boundary: The decision that defines the architecture, and the two that follow from it directly.

  • ADR-01 · The step attempt, not the run, is the unit of durability
  • ADR-02 · The trigger event log is the platform's own system of record for what happened
  • ADR-03 · Acceptance and execution are separated by the log and fail independently

Effects in other people's systems: What a retry means when the thing being retried created an invoice somewhere we do not control.

  • ADR-04 · Every action declares a replay-safety class, and publication is gated on it
  • ADR-05 · A retry reuses the effect key; a replay mints a new logical attempt
  • ADR-06 · An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically

Living inside somebody else's quota: Rate limits we did not set, cannot negotiate, and are punished for ignoring.

  • ADR-07 · Every outbound call leaves through a quota governor; no runner calls a provider directly
  • ADR-08 · A rate limit is a park with a release time, not a failure
  • ADR-09 · Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation

Triggers from systems that cannot push: How the platform finds out that something happened, and how it notices that it has stopped finding out.

  • ADR-10 · Push where available, poll everywhere else, and own the subscription lifecycle as a component
  • ADR-11 · A cursor advances only after durable commit, and absence of data is itself monitored

Credentials held on somebody's behalf: Custody of millions of delegated grants into other companies' products.

  • ADR-12 · Credential custody is a separate account; runners receive single-connection, single-run tokens
  • ADR-13 · Refresh is single-flighted per connection, and credential death is terminal rather than retryable

Tenancy, fairness and the long tail: Twelve million automations, most of which do nothing most of the time.

  • ADR-14 · Fair scheduling on shared capacity, with runaway detection as part of the same decision
  • ADR-15 · Near-zero idle cost is an architectural constraint, not an optimisation

Legibility: The author is not an engineer, and that is an architectural constraint rather than a UX preference.

  • ADR-16 · Every failure carries a classified, author-facing cause and a repair path

Technology by capability

Amazon Web Services was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases self-hosted open source carries 21 of 47 documented stacks and Microsoft Azure 13, with AWS on 9 — and Google Cloud, lower still at 7, took three of the seven most recent documents. Reaching for the same cloud each time teaches a service catalogue rather than architecture. The second is fit, in a weak sense that is worth stating honestly: this topic belongs to no cloud at all. Its hard parts — custody of delegated third-party credentials, a governor for quotas nobody hands you, a stable published egress identity, and a sandbox for partner code — are things every cloud leaves you to build. What AWS supplies well is the shape underneath: a very large fleet of short-lived, bursty, egress-heavy workers behind a fixed NAT range, with a partitioned durable log beside a high-write key-value store. Everything in ask.md is written vendor-neutrally — 'durable event log', not a product name — and the table below is where the neutral capability meets a specific service, with what would be used instead on another stack.

Capability Choice Origin Credible alternative Why this one Record
Trigger event log — the platform's system of record for what happened Amazon MSK, partitioned by connection, 30 days retained as the replay window with tiered storage Amazon Web Services Confluent Cloud or self-managed Kafka; Azure Event Hubs with Capture; Google Pub/Sub with a replay subscription Per-connection ordering without global ordering, and a retention window that is a product promise about how far back a repair can reach ADR-02
Step ledger — run state, written twice per step DynamoDB, keyed (run_id, step_id, attempt), three-AZ synchronous replication, on-demand capacity Amazon Web Services Cloud Spanner or Bigtable; Azure Cosmos DB; CockroachDB or Cassandra self-hosted The hottest write path in the platform is a keyed append with no cross-row transaction; a relational store would buy consistency the access pattern does not need at a cost it cannot absorb ADR-01
Push ingestion endpoints ALB in front of Fargate services, per-connection signature verification, commit before acknowledgement Amazon Web Services Azure Front Door with Container Apps; Google Cloud Load Balancing with Cloud Run; Envoy on Kubernetes A stateless tier whose only obligation is to authenticate and durably commit inside the provider's delivery timeout ADR-03
Trigger cursors and subscription state DynamoDB, one small hot item per connection, advanced only after durable commit Amazon Web Services Azure Cosmos DB; Cloud Bigtable; Redis with AOF persistence, accepting the durability trade Two million connections of small, very hot, very frequently updated state — and the one store whose loss is not recoverable from the two authoritative ones ADR-11
Run admission and queue classes SQS, four queues by latency class, with per-workspace concurrency leases held in DynamoDB Amazon Web Services Azure Service Bus; Google Cloud Tasks; NATS JetStream or RabbitMQ self-hosted Separated queues make the fairness guarantee structural: a bulk import cannot share a queue with a two-step automation even by accident ADR-14
Step runners Fargate tasks, autoscaled on queue depth and age, holding no run state between steps Amazon Web Services Azure Container Apps; Cloud Run; Kubernetes with KEDA Serverless containers keep idle cost near zero across twelve million mostly-dormant automations, which rules out any per-automation reservation ADR-15
Quota governor and egress Fargate service owning per-connection and per-connector leases, behind NAT gateways with a published Elastic IP range Amazon Web Services Azure NAT Gateway with Container Apps; Cloud NAT with Cloud Run; Envoy with a rate-limit service on Kubernetes The quota decision, the token exchange and the egress identity belong at one boundary, and the address range is a customer-visible commitment ADR-07
Credential custody A separate AWS account running the custody service, with per-workspace KMS keys and envelope-encrypted material Amazon Web Services Azure Key Vault in a separate subscription; Cloud KMS in a separate project; HashiCorp Vault or OpenBao with a transit backend The blast-radius boundary is an account boundary, not a service boundary: a compromised task role in the platform account must not reach key material ADR-12
Connector sandbox AWS Lambda per invocation, no credentials in scope beyond the connection, declared egress hosts, hard CPU and memory ceilings Amazon Web Services Azure Functions; Cloud Functions; gVisor or Firecracker microVMs on Kubernetes; a WASM runtime for declarative-only connectors Partner code must not share a process, a filesystem or a network namespace with the run it serves ADR-04
Definition registry and tenancy Aurora PostgreSQL, immutable automation versions with a published-version pointer Amazon Web Services Azure Database for PostgreSQL; Cloud SQL or Spanner; PostgreSQL with Patroni self-hosted Publication is strongly consistent, low-volume and relational; the hot path never reads it without a cache ADR-05
Connector catalogue DynamoDB for manifests and version pointers, S3 for content-addressed artefacts Amazon Web Services Azure Cosmos DB with Blob Storage; Firestore with Cloud Storage; an OCI registry with MinIO Hot reads of varied-shape documents, with immutable artefacts addressed by digest so a pinned version is genuinely the same bytes ADR-04
Run history and analytics projections DynamoDB run index with a TTL by retention plan, S3 and Athena for cold history and cost analysis Amazon Web Services Azure Data Explorer; BigQuery; ClickHouse self-hosted Both are rebuildable from the ledger, so their schema can change without a migration and their loss is a degradation rather than an incident ADR-16

The decisions, and the alternatives that lost

The durability boundary

The decision that defines the architecture, and the two that follow from it directly.

ADR-01 · The step attempt, not the run, is the unit of durability

Status: Accepted · Shown on views: 12, 14, 07, 11

When a worker dies halfway through a five-step automation that has already created an invoice and sent an email, what happens next — and who pays for it?

Context. A run in this platform is a sequence of effects in systems the platform does not own and cannot roll back. If run state lives in the worker's memory, the only recovery from a lost worker is to re-run from the beginning, which means re-performing every step that already succeeded. For a pipeline of pure computation that is a performance problem. Here it is a wrong answer in a customer's accounting system. The same argument applies to every other interruption: a rate limit that needs waiting out, a credential that died mid-run, a provider returning 503. Each of them needs the run to stop and later continue from exactly where it was, and none of them can be served by a design whose only verb is 'start over'. At 12,000 step attempts a second the cost of externalising that state is not theoretical either, which is why this is a decision rather than an obvious default.

Decision. A run is a resumable state machine whose complete state lives in an append-only step ledger outside the worker. Every step attempt writes its intent to the ledger before the outbound call and its outcome after, carrying the run, step, attempt number, resolved inputs, effect key, outcome and the provider's response identifier where one is returned. Run status is a projection of the ledger, never an independently updated field. A worker holds only a lease; losing it releases the run for another worker to resume at the last committed step. No step may be attempted without a stable effect key (ADR-05).

How it is realised on AWS. DynamoDB holds the ledger, keyed (run_id, step_id, attempt), replicated synchronously across three availability zones, with on-demand capacity and sharding by run id. Step runners are Fargate tasks that hold no state between steps; a run lease is a short-TTL item whose expiry makes the run eligible for another runner. The intent write and the outcome write are separate items rather than an update, so the ledger is genuinely append-only and the ambiguous case — intent present, outcome absent — is directly observable rather than inferred. Run status in the author-facing run index is materialised from the ledger asynchronously and can be rebuilt from it in full.

Option Verdict Reasoning
Step-level durability in an external append-only ledger Chosen A park, a retry, a pause and a resume all become the same operation on the same structure, and no completed effect is ever repeated
Run-level durability: the worker holds the run and retries it from the start Rejected Dramatically simpler and needs no ledger, but re-runs step one to retry step four — unacceptable when step one created an invoice
Checkpoint only at explicit author-declared boundaries Rejected Pushes a correctness decision onto a non-engineer who has no way to reason about it, and leaves the default unsafe
A single mutable run row updated in place per step Rejected Loses the intent-without-outcome signal that makes an ambiguous effect detectable, which is the case the whole design exists to handle
An embedded workflow engine holding state in its own store Rejected A reasonable alternative that moves the same decision behind a dependency; rejected because the effect-key semantics in ADR-05 are specific enough that owning the ledger is simpler than bending somebody else's

What it buys

  • A worker loss, a rate-limit park, a credential pause and a transient retry are one mechanism — moving a run between states — rather than four special cases with four recovery paths
  • The ambiguous outcome is directly observable as an intent with no matching outcome, which is what makes ADR-06 implementable at all
  • Run history, the repair inbox and author-initiated replay are all reads of one structure, so they cannot disagree with each other
  • The run index can be dropped and rebuilt, which turns a schema change to author-facing history into a routine operation

What it costs

  • Two synchronously replicated writes per step attempt, which at 12,000 steps/second is the platform's hottest path and its primary scaling constraint
  • A priced component of every step: the ledger write is held to ≤ 18% of per-step cost, and that ceiling is a design constraint rather than a target
  • Latency added to every step by the intent write before the call, paid even on the overwhelming majority of steps that succeed first time
  • A lease mechanism with its own failure modes — a slow worker whose lease expires while its call is in flight is a duplicate risk that ADR-05 has to absorb

Choose differently when. Choose run-level retry instead when the effects are genuinely idempotent end to end, when runs are short enough that re-running from the start is cheap, or when the platform owns the systems the steps touch and can roll them back. A platform automating only its own internal services has all three properties, which is exactly why the adjacent distributed-workflow-orchestration-platform use case can make a different choice.

Why it holds up over time. Externalising long-running process state into a durable log is the same decision a transaction log, a saga and a workflow engine each make, and it has survived every generation of the technology underneath it. The store will be replaced; the property that a worker may die at any instant without repeating a completed effect follows from the effects being irreversible and somebody else's, which will not change.

Lesson. Decide what the unit of durability is before anything else, and derive the rest from it. If the thing you are retrying is irreversible and lives in a system you do not own, the unit cannot be the whole job.

ADR-02 · The trigger event log is the platform's own system of record for what happened

Status: Accepted · Shown on views: 09, 10, 13

When an author asks us to re-run last Tuesday's missed work, what do we read — our copy of the event, or the provider's current state?

Context. The providers are the system of record for what is true. They are not a reliable system of record for what happened: a webhook delivery is usually not retrievable afterwards, a polled change may have been superseded, and a deleted record simply is not there. If the platform keeps no copy, then a replay can only re-read the provider's current state, which means a repair performed a week after a credential died may run against data that has changed or vanished — or may be impossible because the credential is still dead, or because the provider is down. Keeping a copy makes replay exact and independent, and makes the platform the custodian of a month of its customers' data from every product they have connected.

Decision. Every accepted trigger event is durably committed to a partitioned, append-only event log before the provider is acknowledged, and retained for a declared replay window of 30 days. The log is the only thing a replay reads. Retention is therefore a recovery commitment — a promise about how far back a repair can reach — and not a storage preference.

How it is realised on AWS. Amazon MSK, partitioned by connection so per-connection ordering holds without requiring global ordering, with 30 days of retention using tiered storage so the window is affordable. Payload bodies above a size threshold are written to S3 and the log carries a reference, which keeps the log's partitions small and puts the residency and classification rules on object storage where they are easier to enforce. Admission reads the log continuously and deduplicates into at most one run per (automation version, event id).

Option Verdict Reasoning
Retain the payload for a declared replay window; replay reads the log Chosen Replay is exact and independent of the provider's availability, credential state and current data
Retain only metadata and re-fetch the payload from the provider on replay Rejected Holds almost nothing and materially reduces the breach surface, but makes replay silently run against changed data and impossible when the credential is dead — which is the most common reason a replay is needed
Retain nothing; a missed trigger is simply missed Rejected Coherent and honest, and would remove a whole class of risk, but it deletes the repair journey the product is largely sold on
Let the workspace choose the window, including zero Chosen Adopted as a per-workspace data-classification setting alongside the default, with the repair path's reduced capability stated plainly rather than discovered

What it buys

  • A repair performed days later runs against exactly what arrived, not against what the provider happens to hold now
  • The accept path owes the provider durability and nothing else, which is what makes a 250 ms acknowledgement achievable
  • Admission, dedup and fair scheduling can all be rebuilt or re-run from the log without involving any provider

What it costs

  • Thirty days of customer data from every connected product concentrated in one platform — a residency, retention and breach surface far larger than the platform's own records
  • Retention is now a throughput obligation too: the log must be readable fast enough that a bulk replay does not starve live admission
  • A per-workspace no-retention setting creates a second, weaker repair path that has to be explained rather than hidden

Choose differently when. Choose metadata-plus-re-fetch when the payloads are highly sensitive, when providers offer a stable read-by-id that survives the retention period, and when replays are typically attempted within minutes rather than days. Choose to retain nothing when the product makes no repair promise at all.

Why it holds up over time. Separating 'what is true' from 'what happened' is a distinction every event-driven system eventually discovers, usually after trying to reconstruct history from current state and failing. The storage will change; the need for an independent record of the event, owned by the party that must act on it, will not.

Lesson. If you promise to replay something, you must own a copy of it. A replay that reads the source is not a replay, it is a new run wearing the same name.

ADR-03 · Acceptance and execution are separated by the log and fail independently

Status: Accepted · Shown on views: 02, 12, 13

The execution plane is down. Does the platform tell the provider to go away, or take the event anyway?

Context. Provider push deliveries have short timeouts and limited, inconsistent retry behaviour; many providers retry two or three times and then drop the event permanently. A rejected delivery is therefore usually gone for good, and no amount of later recovery brings it back. Execution, by contrast, can be arbitrarily delayed without losing anything, because the work is still described in the log. The two sides of the platform have fundamentally different costs of failure, and giving them one availability target would either over-engineer execution or under-protect ingest.

Decision. The accept path authenticates the delivery, commits it durably to the event log, and acknowledges — and does nothing else. No binding, no credential resolution, no quota check and no execution happens before the acknowledgement. Ingest and execution carry separate availability targets (≥ 99.99% and ≥ 99.95%), scale independently, and a total execution outage loses no trigger.

How it is realised on AWS. ALB in front of stateless Fargate push services; the only work before acknowledgement is signature verification against the per-connection secret and the MSK produce with acknowledgement from all in-sync replicas. Admission runs as a separate consumer group entirely, so its failure, backlog or redeployment is invisible to the provider. Poll workers follow the same contract: read, commit, then advance the cursor (ADR-11).

Option Verdict Reasoning
Authenticate, commit, acknowledge — nothing else on the accept path Chosen The only design in which ingest availability is genuinely independent of everything downstream
Execute the first step synchronously before acknowledging Rejected Gives the author a faster perceived start and ties the acknowledgement to credential and provider availability, which is precisely the coupling that loses events
Validate the payload against the automation's bindings before acknowledging Rejected Catches author errors a few seconds earlier at the cost of making the accept path depend on the definition registry; the same error is caught at publish time instead (ADR-16)
Reject deliveries when the execution backlog is beyond a threshold Rejected Protects the platform by discarding exactly the events it exists to not lose; back-pressure is declared through widened latency budgets instead

What it buys

  • A 250 ms p99 acknowledgement fits comfortably inside every common provider delivery timeout, while a run may legitimately take minutes or park for hours
  • Execution can be redeployed, scaled, or fail entirely without a single trigger being lost
  • Ingest is stateless apart from cursors, so it scales horizontally against burst without coordination

What it costs

  • The author's first feedback moves from the acknowledgement to the run history, so perceived latency has to be managed explicitly as a product surface
  • An automation whose definition is broken still accepts and logs events, producing runs that fail immediately and need a classified cause rather than a rejection
  • Two availability targets mean two operational postures, and the ingest tier must be genuinely over-provisioned rather than merely autoscaled

Choose differently when. Collapse the two when deliveries are reliably retried by the source for hours, or when the platform owns the source. A system reading from its own durable queue has no reason to separate acceptance from execution, because nothing is lost by refusing.

Why it holds up over time. The principle that the acceptance of work and the performance of work have different costs of failure, and therefore different availability targets, outlives any particular queue. It is the same reason a payment gateway acknowledges before settling.

Lesson. Find the step where failure is irreversible — usually the one where an external party gives up — and make it depend on as little as possible.

Effects in other people's systems

What a retry means when the thing being retried created an invoice somewhere we do not control.

ADR-04 · Every action declares a replay-safety class, and publication is gated on it

Status: Accepted · Shown on views: 08, 14, 16

Can this step be retried safely? And who is supposed to know the answer — the platform, the connector developer, or the author?

Context. The platform's execution guarantees cannot be stronger than what the providers support, and support varies wildly across 40,000 actions. Some accept a caller-supplied idempotency key. Some offer a read-back by natural key that reliably shows whether the effect landed. Many — notably 'post this message' and 'send this email', among the most-used actions in the product — offer neither, and after a timeout the platform genuinely cannot know what happened. A single platform-wide retry policy must therefore be wrong for most of the catalogue: safe enough for one class, and either duplicating or silently dropping effects for another.

Decision. Every action carries a declared replay-safety class — idempotent, checkable or unsafe — as a versioned field of the connector manifest. The class is a hard release gate: an action cannot be published without it. Execution semantics are derived from the class rather than configured separately, and the author-facing behaviour of the unsafe class is stated in the product.

How it is realised on AWS. The class lives in the connector manifest in DynamoDB alongside the schemas and rate-limit profile, versioned with the connector and pinned by each automation to a compatible range. The release pipeline rejects a manifest without it, and contract tests run against provider sandboxes rather than mocks, because the failure being guarded against is the provider's behaviour, which a mock cannot reveal. The step runner reads the class at bind time and selects the ambiguity-resolution path in ADR-06.

Option Verdict Reasoning
A declared, versioned per-action class, gated at release Chosen Puts the answer where the knowledge is, and makes it reviewable, testable and pinnable
One platform-wide retry policy Rejected Simple and wrong for most of the catalogue: necessarily either duplicating or dropping, depending on which way it is tuned
Infer the class at runtime from provider responses Rejected Attractive and unsound: the inference is only testable by causing the ambiguity it is meant to resolve, in a customer's system
Let the author choose per step Rejected Asks a non-engineer a question they have no basis to answer, and the wrong answer is invisible until an invoice is duplicated
Let the author choose only the unsafe class's behaviour Chosen Adopted in ADR-06: the author decides what to do about an ambiguity, never whether the ambiguity exists

What it buys

  • The platform can state a different, honest guarantee per action instead of one promise that is false for most of them
  • The interface catalogue becomes organised by the property that actually matters, which makes the unsafe surface countable and reviewable
  • Pinning a connector version pins the guarantee, so a provider change cannot silently alter execution semantics for a published automation

What it costs

  • The class is self-declared by connector developers whose incentive is to ship, and a wrong declaration is indistinguishable from a right one until an effect is duplicated
  • A per-action field across 40,000 actions is a large surface to review and to keep honest as providers change
  • The product must explain a distinction users did not ask for, at the moment they are least interested in it

Choose differently when. Drop the classification when every integration target is under your control and can be made idempotent by fiat. An internal platform can mandate idempotency keys across its own services and needs none of this.

Why it holds up over time. For as long as a platform integrates systems it does not control, something must carry the answer to 'is a retry safe here'. The mechanism may one day be discovered rather than declared, but the dependency on the answer is permanent.

Lesson. When a guarantee depends on a third party's capability, make the capability an explicit, versioned, testable field — not an assumption spread through the code that uses it.

ADR-05 · A retry reuses the effect key; a replay mints a new logical attempt

Status: Accepted · Shown on views: 11, 12, 14

The author presses a button marked 'run it again'. Did they mean 'the first one may not have worked' or 'do it a second time on purpose'?

Context. Those two intentions produce opposite correct behaviours, and a platform that conflates them will sometimes duplicate an invoice and sometimes fail to send one. The distinction cannot live in a flag read at call time, because by then the two paths share the same code and the same provider call; it has to be built into how the key is derived, so that the difference is structural and cannot be got wrong by a caller.

Decision. Each step attempt carries a deterministic effect key derived from (run id, step id, logical attempt). A retry is a new physical attempt at the same logical attempt and therefore reuses the key, which is passed to the provider as an idempotency key wherever the action's class says it will be honoured. A replay is an author-initiated action that mints a new logical attempt and therefore a new key, is recorded in the ledger as a replay, and requires the duplicate consequence to be stated and confirmed first.

How it is realised on AWS. The key is a hash of the run id, step id and logical attempt number, computed by the runner before the intent write so that the ledger entry and the outbound call carry the same value. The governor forwards it in whatever header the connector manifest declares. The ledger's attempt key distinguishes physical attempts, while the logical attempt is a field on the entry — so 'how many times did we try' and 'how many times did the author ask for this' are separately countable.

Option Verdict Reasoning
Key derived from (run, step, logical attempt); replay increments the logical attempt Chosen Makes the distinction structural: the same code cannot deliver a replay when a retry was meant
A random key per physical attempt Rejected Defeats the entire purpose; every retry becomes a new effect at the provider
A key derived from the payload's content hash Rejected Collides across two legitimately identical actions — two identical messages the author genuinely wanted twice — and silently drops the second
One key per run Rejected Cannot distinguish steps, so a provider honouring it would reject the second step of the same run
Let the connector supply its own key derivation Rejected Flexible, and makes the platform's correctness depend on 8,000 independent implementations of the same subtle rule

What it buys

  • A retry after a timeout, a lease expiry or a park cannot produce a second effect where the provider honours keys
  • 'Try again' and 'do it again' are visibly different in the ledger, so the audit answers which one happened
  • The key is transport-independent and survives a change of provider API or protocol

What it costs

  • The key must be minted before the intent write, so the derivation sits on the hottest path and cannot depend on anything slow
  • It is only useful where the provider honours it — the unsafe class carries the key and gets nothing for it, which has to be explained rather than assumed
  • A lease expiry during an in-flight call produces two physical attempts at the same logical attempt, which is safe only because the key is reused — making lease handling a correctness concern, not just a scheduling one

Choose differently when. A single key per operation is enough when the system performs one effect per job and never needs an intentional repeat. The two-level scheme earns its complexity only where a deliberate re-do is a product feature.

Why it holds up over time. Caller-supplied operation identity is an old idea with new names each decade — message dedup id, idempotency key, request id. What endures is that the caller, not the callee, has to assert what counts as the same operation.

Lesson. When two user intentions demand opposite behaviours, encode the difference in the data structure rather than in a parameter, so no code path can confuse them.

ADR-06 · An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically

Status: Accepted · Shown on views: 14, 21, 12

The request was sent, the connection dropped, and the provider offers no key and no read-back. Do we send it again or not?

Context. This is the one case where the platform genuinely cannot know what happened and cannot find out. Retrying risks a second email to a customer or a second message in a channel. Not retrying risks the thing never happening at all, silently. Both are wrong, and which is less wrong depends entirely on the action and the author's situation — a duplicate invoice is a serious problem, a duplicate internal reminder is not. The platform has no basis to choose, and choosing silently means being wrong without telling anybody.

Decision. For an action classified unsafe, an ambiguous outcome is never retried automatically. The run is parked in a state that names the ambiguity, the author is told exactly what may or may not have happened, and two explicit choices are offered: continue without retrying, accepting a possible gap, or retry and accept a possible duplicate. For checkable actions the platform reads back by natural key before retrying. For idempotent actions it simply retries under the same key.

How it is realised on AWS. The runner detects the ambiguity directly from the ledger's shape — an intent with no matching outcome — and branches on the manifest's class. The parked state carries a typed reason and the provider's last response if any, and surfaces in the repair inbox as its own category rather than mixed with ordinary failures. The run deadline still applies, so a parked-ambiguous run terminates into a named terminal state rather than waiting indefinitely for an author who will never look.

Option Verdict Reasoning
Park and escalate to the author with two explicit choices Chosen The platform refuses to guess on the author's behalf about an effect in their own business
Always retry Rejected Produces duplicate invoices and duplicate customer emails, which is the failure customers escalate hardest and trust least
Never retry Rejected Produces a silent gap, which is worse because nobody finds out — the exact failure mode journey 05 is about
Let the connector developer pick the default Rejected Moves the guess one step away without improving the information available to make it
Refuse to publish automations whose first write is unsafe without acknowledgement Deferred Attractive for high-stakes actions and left to Core Architecture Question 3; it raises a friction cost at authoring time that has not been measured

What it buys

  • The platform never silently creates a duplicate effect in a customer's finance or communications system
  • The ambiguity is named and visible rather than resolved by a default nobody chose
  • The decision is made by the only party with the business context to make it

What it costs

  • A parked run waits for a human, and many authors will not look — the run deadline then terminates it, which is itself a form of the silent gap this decision tried to avoid
  • The product must explain a genuinely hard concept at a bad moment, and 'we do not know if your email sent' is a difficult sentence
  • Parked-ambiguous runs accumulate and need their own inventory, alerting and expiry policy

Choose differently when. Pick a default and apply it when the actions are low-stakes and homogeneous — an internal notifier can always retry, and nobody will mind. The escalation is worth its friction only where the effects are irreversible and matter commercially.

Why it holds up over time. Distributed systems will not stop producing ambiguous outcomes, and no protocol removes the case where the acknowledgement is lost and the operation cannot be queried. The lasting decision is whose judgement resolves it.

Lesson. When a system cannot know the answer and the consequences of guessing fall on someone else, surface the choice rather than picking a default and calling it a policy.

Living inside somebody else's quota

Rate limits we did not set, cannot negotiate, and are punished for ignoring.

ADR-07 · Every outbound call leaves through a quota governor; no runner calls a provider directly

Status: Accepted · Shown on views: 07, 12, 19

Twelve thousand runners want to call the same provider at once. Who decides whether they may?

Context. Providers meter access per end-user account, per connector application credential, per region and per source address, and the penalties for overshoot escalate from a 429 to a suspended application credential that breaks every workspace at once. A decision made independently by each runner cannot respect a limit that is shared across runners, and the shared cases are the dangerous ones: a per-application limit is one scarce resource that every tenant is drawing on simultaneously. The same boundary is also the natural place to exchange a short-lived credential, because it is the last point before the call leaves the platform.

Decision. Every outbound third-party call passes a governance point that owns the quota decision; no step worker may call a provider on its own judgement. The governor enforces limits at each level the provider actually imposes them, honours provider backoff headers as authoritative over its own schedule, exchanges the short-lived single-connection token with custody, opens per-provider circuit breakers, and exits through the published egress range.

How it is realised on AWS. A Fargate service holding per-connection and per-connector leases in DynamoDB, in front of NAT gateways with a published Elastic IP range. Leases are checked out in blocks sized to the connection's observed rate so the common case is a local decision, with central arbitration reserved for per-application limits where tenants genuinely contend. Circuit breaker state is per provider and shared across the fleet, so one runner discovering an outage protects the rest.

Option Verdict Reasoning
A governance point owning the decision, with leases checked out in blocks Chosen Correct for shared limits, and degrades to local decisions with bounded overshoot when the governor is slow
Per-worker local budgets with periodic reconciliation Rejected Fastest and overshoots exactly when traffic is burstiest, which is when a provider is least forgiving
A central broker consulted synchronously on every call Rejected Correct and puts a hard synchronous dependency plus a network hop on the hottest path in the platform
Rely on the provider's 429 and back off reactively Rejected Makes the provider's rate limiter our rate limiter, which works until the provider responds by suspending the application credential

What it buys

  • A limit shared across tenants can actually be respected and fairly apportioned, which no local scheme achieves
  • Credential exchange, egress identity and the quota decision are one boundary, which is what keeps long-lived credentials out of the execution plane
  • Provider health becomes a first-class platform signal rather than something each runner rediscovers

What it costs

  • The governor is on the path of every outbound call and is the platform's hottest single dependency
  • Leases add a tuning problem: too small and they are a central broker with extra steps, too large and they overshoot during a burst
  • A cross-account token exchange per run adds latency, mitigated by caching for the run's duration at the cost of widening the revocation window

Choose differently when. Local budgets are sufficient when every limit is per end-user account, traffic is smooth, and the penalty for a modest overshoot is only a 429. Central arbitration earns its cost when tenants share one scarce credential.

Why it holds up over time. Providers will always meter, and the penalty for overshoot will always exceed the cost of pacing. Putting that decision in one place rather than in every caller survives any change of rate-limiting algorithm.

Lesson. A constraint that is shared cannot be enforced by parties acting independently. Put the decision where the constraint is, even when that means a hop on the hot path.

ADR-08 · A rate limit is a park with a release time, not a failure

Status: Accepted · Shown on views: 12, 21, 17

A provider says 'not now, try in ninety seconds'. Is that an error?

Context. At this scale a 429 is the ordinary case, not an incident — the platform makes billions of calls a month against quotas it did not set. Treating it as a failure has three bad consequences: it consumes the step's retry budget on something that is not a fault, it occupies a worker for the duration of a backoff, and it eventually exhausts the attempts and fails a run that was never actually broken. It also teaches the platform to treat the provider's own backoff signal as advice rather than instruction, which is how an application credential gets suspended.

Decision. A rate-limit response moves the run to a waiting state with a release time derived from the provider's own backoff header, which is authoritative in preference to the platform's schedule. A park consumes no worker and does not spend the step's retry budget. The connection's pacing adapts in response, so repeated parking feeds back into the lease size rather than simply repeating.

How it is realised on AWS. The governor returns a typed park outcome with a release timestamp; the runner writes it to the ledger and releases its lease. A dedicated parked-release queue class holds the run until its time, so parked work is isolated from interactive runs and cannot distort their queue-age signal. The platform's own SLO is that ≤ 0.1% of outbound calls receive a 429 at all — the governor's job is to make parking rare, not to make it cheap.

Option Verdict Reasoning
Park with a provider-derived release time, outside the retry budget Chosen Treats the ordinary case as ordinary, and keeps the retry budget for actual faults
Treat 429 as a retryable error with exponential backoff Rejected Spends the retry budget on capacity rather than faults, and holds a worker through every backoff
Fail the run and let the author retry Rejected Converts a routine, self-resolving condition into manual work for someone who cannot influence it
Use our own backoff schedule and ignore Retry-After Rejected Substitutes our guess for the provider's statement about its own capacity, which is both rude and usually wrong

What it buys

  • Throughput recovers at the provider's stated time rather than at an arbitrary multiple of it
  • The retry budget stays meaningful as a signal about faults, which keeps the failure taxonomy honest
  • Parked runs occupy no compute, so a provider-wide slowdown costs storage and patience rather than capacity

What it costs

  • A parked-release queue class with its own scheduling, inventory and alerting
  • Parked runs can accumulate into a backlog that fires in a burst at release, which needs its own rate-limited release path
  • The author experiences a slow automation with no error, which must be explained in the run history as 'running slowly' rather than left blank

Choose differently when. Treat rate limits as errors when they are genuinely exceptional and indicate misconfiguration rather than capacity — a well-provisioned internal API where a 429 means somebody has a bug.

Why it holds up over time. Backpressure from a dependency is information, not failure. Systems that learn to wait on it outlive systems that learn to retry through it, and that has been true of every protocol that has ever had a busy signal.

Lesson. Distinguish 'this is broken' from 'not right now'. Budgets, alerts and retries that conflate them will spend themselves on the wrong thing.

ADR-09 · Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation

Status: Accepted · Shown on views: 15, 19, 07

A customer's security team wants to allowlist us. What address do we give them, and what happens when a neighbour abuses it?

Context. Enterprise customers and many providers restrict inbound API access by source address, so a platform that cannot name its egress addresses cannot be adopted by them at all. Publishing a range makes it a commitment: it cannot change without notice, which constrains every future networking decision. It also makes the range a shared identity — providers rate-limit, throttle and block by source address, so a single runaway automation can degrade or sever access for every workspace behind it.

Decision. The platform egresses through NAT gateways with a published, stable address range that customers and providers may allowlist, and treats changes to it as a product change requiring notice. Because the range is shared, runaway detection, per-connection and per-connector quota leases, and the ability to move a workspace onto a separate egress range are part of the same decision rather than separate features.

How it is realised on AWS. NAT gateways with Elastic IPs per availability zone, published and versioned as documentation. Every outbound request carries a connector-identifying user agent and a correlation identifier a provider can quote in a support ticket. Residency-pinned workspaces egress from their own region's range. Runaway detection throttles an automation within 60 s of exceeding 10x its 7-day baseline, before a provider notices rather than after.

Option Verdict Reasoning
A published stable NAT range, with runaway controls as part of the same decision Chosen Enterprise adoption requires it, and the shared-reputation risk it creates has to be mitigated in the same breath
Ephemeral egress addresses Rejected Operationally simplest and excludes every customer whose security policy requires an allowlist
A dedicated egress address per workspace Rejected Removes the shared-reputation risk entirely and does not scale to 2.5 M workspaces at any sane cost
Dedicated egress for workspaces that require or earn it Chosen Adopted as both a plan feature and an incident response, on the same code path

What it buys

  • Enterprise customers can allowlist the platform, which is a precondition for a large part of the market
  • A provider investigating traffic can identify the caller and quote a correlation id back, turning an opaque block into a conversation
  • Residency pinning extends naturally, because the egress identity is already per-region

What it costs

  • The range becomes immovable without a notice period, constraining future networking changes
  • One workspace's abuse is every workspace's problem, which raises the stakes on runaway detection considerably
  • Per-workspace egress as an escape hatch needs to exist before it is needed, not after

Choose differently when. Skip the published range when no customer requires allowlisting and providers do not restrict by address — a consumer-only product can egress from anywhere and avoid the shared-reputation problem entirely.

Why it holds up over time. Network identity as a trust signal has survived every change of protocol and will outlast the current generation of address-based controls, because the underlying need — naming who is calling — does not go away.

Lesson. A shared identity is a shared liability. If you publish one, you must also own the controls that stop one tenant spending the reputation of all of them.

Triggers from systems that cannot push

How the platform finds out that something happened, and how it notices that it has stopped finding out.

ADR-10 · Push where available, poll everywhere else, and own the subscription lifecycle as a component

Status: Accepted · Shown on views: 13, 08, 17

Four fifths of the catalogue will not tell us when something changes. How does the platform find out, and how does it notice that it has stopped finding out?

Context. Provider push is cheap, fast and available for roughly a fifth of the catalogue, and it brings subscription lifecycle, signature verification, duplicate deliveries and — most damagingly — silent subscription death, where a subscription expires or is revoked and the platform simply stops receiving events with no error anywhere. Polling works everywhere, costs linearly in connections, and at 2,000,000 connections on an average five-minute interval amounts to roughly 6,700 provider polls a second that somebody has to pay for and that providers have to absorb.

Decision. Both, treated as architecturally different rather than hidden behind one abstraction. Push is preferred wherever the provider supports it, and the platform owns the subscription lifecycle as a component: create, renew before expiry, detect silent death, re-create. Polling adapts its interval per connection to the provider's quota profile, the plan's entitlement and the observed change rate, with declared jitter, and the initial poll establishes a watermark rather than emitting an account's entire history.

How it is realised on AWS. A subscription manager tracks every push subscription's renewal deadline and renews ahead of it; a liveness expectation per connection raises a quiet-period alarm when a source has emitted nothing for longer than its declared quiet period, which is wired off the event log so it watches for the absence of data rather than the health of the poller. Poll workers are Fargate tasks reading cursors from DynamoDB with per-connection jitter, backing off for quiet connections and batching provider queries where the API allows.

Option Verdict Reasoning
Hybrid, with the subscription lifecycle owned explicitly Chosen The only design that covers the catalogue while giving the fifth that can push the latency it deserves
Poll everything, for uniformity Rejected One mechanism and one failure taxonomy, at the cost of latency on the connectors where latency is most visible and a much larger polling bill
Push only, and omit connectors that cannot push Rejected Would cut the catalogue — the growth ceiling — by roughly four fifths
Treat polling as a fallback inside one trigger abstraction Rejected Hides two genuinely different failure taxonomies behind one interface, which is how a dead subscription gets mistaken for a quiet account

What it buys

  • Push connectors get seconds of latency and polled connectors get a declared, plan-appropriate interval, instead of both getting the worse of the two
  • A subscription that expires is detected and re-created rather than becoming an invisible outage
  • Jitter prevents the platform from aligning a million connections on the minute boundary against one provider

What it costs

  • Two freshness models, two dedup mechanisms and two failure taxonomies behind one product promise
  • Polling infrastructure is assumed at ≤ 25% of total platform compute cost, attributed to the connectors that lack push support
  • Subscription lifecycle is ongoing work that scales with the catalogue and with provider API churn

Choose differently when. Poll everything when the catalogue is small and latency expectations are loose; push only when every integration target is modern and under contract. The hybrid is the price of a long tail.

Why it holds up over time. The split between systems that notify and systems that must be asked has outlived several generations of integration technology and shows no sign of closing, because it reflects the providers' priorities rather than a technical limitation.

Lesson. Do not unify two mechanisms whose failure modes differ. One abstraction over both hides exactly the distinction operators need when something stops working.

ADR-11 · A cursor advances only after durable commit, and absence of data is itself monitored

Status: Accepted · Shown on views: 13, 10, 17

What is the smallest piece of state in this platform, and what happens when we lose it?

Context. A polled connection's cursor — a watermark, an updated-at bound or a page token — is a few bytes per connection. It is also the only state in the architecture whose loss cannot be recovered from the two authoritative stores: the event log knows what arrived, and the ledger knows what was done, but neither knows what the provider would have returned. Lose a cursor forward and there is a permanent silent gap; lose it backward and there is a flood of duplicates. Worse, the ordinary failure here produces no error at all — a stuck cursor and a healthy quiet account look identical.

Decision. A cursor advances only after the events it covers are durably committed to the event log, never before. The cursor and subscription store is treated as a distinct data zone with its own durability posture rather than grouped with other operational state. Every connection carries a liveness expectation, and an unexpectedly quiet source raises an alarm — the platform monitors the absence of data, not only the health of the components that fetch it.

How it is realised on AWS. DynamoDB, one small item per connection, written after the MSK produce is acknowledged by all in-sync replicas. The quiet-period expectation is derived from the connection's own observed history rather than a global constant, so a genuinely low-traffic connection does not alarm. The alarm is evaluated off the event log, so it fires whether the cause is a dead subscription, a stuck cursor, a paused automation or a provider that has quietly stopped returning results.

Option Verdict Reasoning
Advance after commit, with a per-connection liveness expectation Chosen Guarantees at-least-once at the cost of occasional duplicates, which ADR-05's effect key absorbs
Advance before commit Rejected A one-line difference that produces a permanent, silent, unrecoverable gap — the worst failure mode available to this platform
Make the cursor store as durable as the event log Deferred Removes the residual risk and roughly doubles the cost of the hottest small store; left open as the second half of Core Architecture Question 1
Reconstruct cursors from the event log after a loss Rejected Recovers the last event seen but not the provider's pagination position, so it narrows the gap without closing it

What it buys

  • A crash between read and commit costs a duplicate, which the effect key makes harmless, rather than a gap, which nothing makes harmless
  • A stuck cursor, a dead subscription and a silently broken provider all surface through one alarm that watches for silence
  • The riskiest store in the architecture is named as such, which is the precondition for anyone treating it carefully

What it costs

  • Duplicates are now routine on the poll path and the dedup window has to be sized for them
  • A per-connection liveness baseline is state about state, and it has to be learned rather than configured across two million connections
  • The residual risk is real and unresolved: a cursor store failure still produces duplicates or a gap, and the choice between them is not yet made

Choose differently when. Advancing before commit is defensible only when the source can be fully re-read cheaply and idempotently at any time, which makes the cursor an optimisation rather than state.

Why it holds up over time. Checkpoint-after-commit is the same rule as a consumer offset in a log, a replication position, or a file-read watermark. It will outlive this platform because it follows from the ordering of durability, not from any technology.

Lesson. Find the smallest piece of state whose loss cannot be reconstructed from anything else, and design for that one first. It is rarely the biggest store, and it is usually the one nobody is watching.

Credentials held on somebody's behalf

Custody of millions of delegated grants into other companies' products.

ADR-12 · Credential custody is a separate account; runners receive single-connection, single-run tokens

Status: Accepted · Shown on views: 19, 20, 07

Our execution plane is compromised. How much of our customers' other SaaS estate goes with it?

Context. The platform holds delegated grants into millions of customers' CRMs, mailboxes, file stores and finance systems. Its own data is not the prize; the access is. The execution plane is simultaneously the largest attack surface in the system — it binds untrusted provider responses, runs partner connector code, and makes outbound calls to arbitrary hosts — and the component that needs credentials most often. Those two facts pull in opposite directions, and resolving them by convenience means a single compromised task role reads every refresh token the platform holds.

Decision. Credential custody runs in a separate cloud account with its own key material, reachable only through a narrow exchange interface. No long-lived credential is ever available to a step worker. A worker receives a short-lived token scoped to one connection and one run, exchanged at the egress boundary (ADR-07); the refresh token never leaves custody. Material is encrypted under a per-workspace key, and plaintext never appears in logs, run history, error messages or the ledger.

How it is realised on AWS. A separate AWS account running the custody service on Fargate, with per-workspace KMS keys and envelope-encrypted ciphertext. Cross-account access is by a narrowly scoped role permitting only the exchange operation, never a read. The connection's state and scopes stay in the platform account so the control plane can query them; only the ciphertext and key reference live in custody. Every exchange is written to an audit record retained seven years independently of the workspace's retention setting.

Option Verdict Reasoning
A separate account, with exchange-only access and run-scoped tokens Chosen The blast-radius boundary matches the account boundary, which is the only boundary an attacker cannot talk their way across with a stolen role
A managed secret store in the same account, read by runners Rejected Far simpler and makes a compromised runner equivalent to a compromise of every customer's connected products
A separate service in the same account Rejected Better than nothing and still inside one IAM blast radius, which is the thing being defended against
Hold no credentials: ask the author to re-authorise each run Rejected Eliminates the risk and the product, which exists precisely to run unattended

What it buys

  • A compromised execution plane yields run-scoped tokens for connections currently executing, not the estate
  • Per-workspace keys make crypto-deletion of one workspace's credentials a key operation rather than a data sweep
  • The audit of credential use is kept outside the workspace's own control, so it cannot be shortened by the party it holds accountable

What it costs

  • A cross-account hop on the path of every step that touches a provider, mitigated by per-run caching at the cost of a wider revocation window
  • Two accounts to operate, deploy and keep in step, with their own failure and permission-drift modes
  • Custody becomes a 99.99% dependency — every step needs it — which is a higher bar than the execution plane it serves

Choose differently when. A same-account secret store is proportionate when credentials reach only systems you already own, so a compromise gains the attacker nothing they did not already have.

Why it holds up over time. Separating custody of a credential from use of a credential is the oldest idea in access control and is independent of the technology implementing either. The boundary will move between accounts, enclaves or HSMs; the separation will not go away.

Lesson. Size the blast radius by what an attacker gains, not by what you store. If the value is access to other people's systems, the boundary has to be one your own compromised code cannot cross.

ADR-13 · Refresh is single-flighted per connection, and credential death is terminal rather than retryable

Status: Accepted · Shown on views: 20, 21, 18

A thousand runs on one connection hit an expired access token in the same second. What happens?

Context. Without coordination the answer is a thousand simultaneous refresh attempts against the provider. Many providers rotate the refresh token on use and invalidate the previous one, so concurrent refreshes race: one succeeds, the rest present a token that has just been invalidated, and the connection dies — a credential death entirely manufactured by the platform. The converse error is equally damaging: treating a genuine revocation as a transient fault and retrying produces a sustained authentication storm against a provider that has already said no, which is a reliable way to have an application credential suspended.

Decision. Access-token refresh is serialised per connection with single-flight semantics: one refresh proceeds, concurrent callers wait for its result. Credential death — revocation, expiry, password change, scope removal — is classified as non-retryable. Affected runs are parked, the automation is paused after a declared grace period, and the connection's owner is notified. The platform never retries an authentication failure as though it were transient.

How it is realised on AWS. Custody holds a per-connection lock for the duration of a refresh and returns the new token to all waiters. Failure classification distinguishes an expired access token (refresh and continue) from a dead grant (terminal), using the provider's error code as declared in the connector manifest rather than inferring from a status code alone. The terminal state feeds the automation health loop: classify, hold, notify, repair, observe.

Option Verdict Reasoning
Single-flight refresh per connection; credential death terminal Chosen Prevents the platform from causing the failure it is trying to recover from
Let each run refresh independently Rejected Causes refresh-token rotation races and converts a routine expiry into a dead connection under load
Refresh proactively on a schedule Chosen Adopted alongside, as an optimisation that reduces how often the single-flight path is contended; it does not remove the need for it
Retry authentication failures with backoff Rejected Produces an authentication storm against a provider that has already declined, and risks suspension of the application credential

What it buys

  • A connection under heavy concurrent load refreshes once, not a thousand times
  • A revoked credential fails fast and visibly instead of degrading slowly behind retries
  • Authentication traffic to providers stays proportional to connections rather than to runs

What it costs

  • A per-connection lock in custody, on the path of every step, with its own contention and timeout behaviour
  • Waiters block on a refresh they did not initiate, adding tail latency at exactly the moment a connection is busiest
  • Distinguishing 'expired' from 'revoked' depends on provider error semantics declared in the manifest, which is another field that can be wrong

Choose differently when. Independent refresh is harmless where providers do not rotate refresh tokens and tolerate concurrent refreshes. The single-flight requirement comes entirely from rotation semantics.

Why it holds up over time. Collapsing concurrent identical work into one operation is a pattern older than the web and applies wherever a shared resource is refreshed under load. Token rotation makes it a correctness requirement rather than an efficiency one, and rotation is becoming more common, not less.

Lesson. When a recovery action has side effects on shared state, concurrency control is part of correctness. An unsynchronised refresh is not a slow path, it is a bug.

Tenancy, fairness and the long tail

Twelve million automations, most of which do nothing most of the time.

ADR-14 · Fair scheduling on shared capacity, with runaway detection as part of the same decision

Status: Accepted · Shown on views: 02, 17, 15

One workspace imports fifty thousand rows. What happens to everybody else's two-step automation?

Context. The workload is extremely uneven: most workspaces run a handful of automations occasionally, and a few run bulk operations that enqueue tens of thousands of runs in a burst. On one shared pool with a single queue, the bulk workspace wins simply by arriving, and every other tenant's latency collapses. Worse, some of those bursts are not legitimate work at all but automations triggering themselves — an automation that writes to the sheet it is watching — which can consume a workspace's quota and a provider's goodwill in minutes.

Decision. Shared capacity with fair scheduling across workspaces and per-workspace concurrency leases, so no workspace holds more than 5% of a shared pool for longer than 60 s while others have queued work. Queue classes are separated by expected latency — interactive, retry, parked-release and bulk — so a backfill never shares a queue with a short automation. Runaway detection throttles an automation within 60 s of exceeding 10x its own 7-day baseline, then pauses it with notice.

How it is realised on AWS. Four SQS queues by latency class, with per-workspace concurrency leases in DynamoDB enforced at admission rather than at the runner, so an over-quota workspace never occupies a worker it will not be allowed to use. Runaway baselines are per automation and learned from its own history, because a rate that is pathological for one automation is normal for another. Dedicated capacity for workspaces that require it runs on the same code path, as a lease pool rather than a separate fleet.

Option Verdict Reasoning
Shared pool, fair scheduling, separated queue classes, runaway detection Chosen Best utilisation with a bounded tail, and keeps idle cost near zero (ADR-15)
One shared pool, first-come-first-served Rejected Simplest and gives the worst tail exactly when it is most visible
Pools per plan tier Rejected Simple isolation that strands capacity and still does not protect tenants from each other within a tier
Dedicated workers per workspace Rejected Ends noisy neighbours and breaks the near-zero idle cost that makes twelve million automations affordable
Dedicated capacity as a plan feature on the same code path Chosen Adopted for the workspaces that require it, without creating a second execution path to maintain

What it buys

  • A bulk import is slowed rather than permitted to starve, and the guarantee is structural because the queues are physically separate
  • A self-triggering automation is caught in a minute, before it spends a provider's goodwill or the workspace's quota
  • Utilisation stays high, which is what keeps per-step cost inside its target

What it costs

  • Fair scheduling is real machinery with its own tuning, and its failure mode is a latency regression that is hard to attribute
  • Per-automation baselines are state that must be learned across twelve million automations and will misfire on genuinely spiky but legitimate work
  • If plan-differentiated latency is ever sold, fairness and the commercial promise are in direct conflict and the architecture has to encode which wins

Choose differently when. First-come-first-served is adequate when the workload is homogeneous and no tenant can enqueue orders of magnitude more than another. Dedicated pools win when tenant isolation is a compliance requirement rather than a performance one.

Why it holds up over time. Fair queueing across tenants on shared infrastructure is as old as time-sharing and has survived every change in what the shared resource is. The specific scheduler will change; the need to stop one tenant monopolising a shared pool will not.

Lesson. Make isolation structural rather than behavioural. Separate queues enforce a guarantee that a shared queue with good intentions cannot.

ADR-15 · Near-zero idle cost is an architectural constraint, not an optimisation

Status: Accepted · Shown on views: 15, 17, 07

Twelve million automations exist. Most of them will not run today. What do they cost?

Context. The business is the long tail: the median workspace has a few automations that fire occasionally, and a large fraction of the twelve million are effectively dormant. If a dormant automation costs anything meaningful — a reserved worker, an open connection, a provisioned queue, a polling slot it does not need — the economics fail at a scale the product must reach to be viable. This is a genuine architectural constraint rather than a cost optimisation, because it rules out whole classes of design at the point where they are chosen.

Decision. An enabled but dormant automation holds no reserved compute, no open connection and no worker slot, at an assumed cost of ≤ $0.004 per automation per month. This rules out per-automation reserved capacity, dedicated workers by default, per-automation long-lived connections and any design where an automation's existence rather than its activity drives cost. Poll cost must be sub-linear in connection count, achieved through interval adaptation, change-rate backoff for quiet connections and batched provider queries.

How it is realised on AWS. Fargate tasks autoscaled on queue depth and age, with nothing allocated per automation. Automations exist as rows in the definition registry and subscriptions or cursors in a small key-value store — bytes, not capacity. Poll scheduling adapts per connection, so a connection that has not changed in a month is polled at its plan's floor rather than its ceiling.

Option Verdict Reasoning
Serverless workers, no per-automation allocation, adaptive polling Chosen The only shape where twelve million mostly-dormant automations are economically viable
A reserved worker or process per active automation Rejected Simplest mental model and fails on cost by orders of magnitude at this scale
Long-running connections per automation Rejected Would reduce latency and makes dormant automations expensive in exactly the way the economics cannot absorb
Archive dormant automations to cold storage Deferred Worth doing if registry cost becomes material; it is not, because a dormant automation is already only bytes

What it buys

  • The long tail is viable, which is the business rather than a technical nicety
  • Capacity tracks actual work, so a quiet weekend costs nothing and a Monday burst scales up
  • No per-automation resource means no per-automation leak, exhaustion or cleanup path

What it costs

  • Cold-start latency on scale-up, paid by whichever runs arrive first after a quiet period
  • A per-task premium over long-running instances, accepted deliberately as the price of the constraint
  • Adaptive polling adds state and complexity to the one tier that would otherwise be simple

Choose differently when. Per-tenant allocation is fine when tenants number in the thousands and each is commercially significant. The constraint comes from the tail's length, not from any dislike of reserved capacity.

Why it holds up over time. Making cost proportional to work rather than to configured entities is the economic principle under every generation of elastic infrastructure, and it will outlast the particular runtime that delivers it.

Lesson. When the business model is a long tail, the cost of an idle entity is an architectural constraint. Decide it before the design, because it eliminates options rather than tuning them.

Legibility

The author is not an engineer, and that is an architectural constraint rather than a UX preference.

ADR-16 · Every failure carries a classified, author-facing cause and a repair path

Status: Accepted · Shown on views: 21, 18, 05

A non-engineer's automation failed. What do they see, and what can they do about it?

Context. The product's premise is that someone who cannot read a stack trace can build and own this. That premise turns failure presentation into an architectural concern rather than a UI one: if the platform's internal states cannot be reduced to a sentence an author can act on, then either the states are wrong or the product is undeliverable. The default alternative — forwarding the provider's error — is worse than useless, because provider errors are written for the developer who called the API, and often describe conditions the author has no way to influence.

Decision. Every failure is classified into a small, stable, author-facing taxonomy, each class carrying a named next action. An unclassified error is never presented as the primary message, and an unclassified share above 2% of author-facing failures is treated as a defect in the taxonomy. A repair inbox groups parked and failed runs by cause with bulk retry, replay, discard and reconnect actions. The classification exists in the execution path, not in the presentation layer.

How it is realised on AWS. The runner writes a typed cause onto the ledger entry at the moment of failure, where the provider response and the step context are both still available; the run index projects it, and the repair inbox groups on it. Provider-specific error codes are mapped to platform classes in the connector manifest, so the mapping is versioned with the connector that knows about it. Redaction of credential material and author-declared sensitive fields happens before anything reaches run history.

Option Verdict Reasoning
A typed taxonomy written at failure time, with a named action per class Chosen The only approach where the classification has access to the context needed to be accurate
Forward the provider's error message Rejected Cheapest and describes conditions in a vocabulary the author cannot act on, often about systems they do not administer
Classify in the presentation layer from stored error text Rejected Brittle string matching, far from the context, and silently wrong whenever a provider rewords a message
Route every failure to human support Rejected Does not scale past a tiny fraction of 1.8 billion monthly step attempts, and is slower than the author fixing it themselves

What it buys

  • An author can usually repair their own automation without support, which is the only workable path at this volume
  • The unclassified rate becomes a measurable quality signal for the taxonomy itself, so gaps surface rather than accumulate
  • Grouping by cause makes bulk repair possible: one reconnection fixes hundreds of parked runs at once

What it costs

  • A mapping from provider errors to platform classes, per connector, which is ongoing work and another field that can be wrong
  • The taxonomy must stay small to be legible and complete enough to cover reality, and those pull against each other continuously
  • Redaction before storage means the raw provider response is not always available for operator debugging, which has its own cost

Choose differently when. Forwarding raw errors is correct when the users are developers who called the API themselves and have the context to interpret them. The taxonomy's cost is justified only by the non-engineer premise.

Why it holds up over time. The requirement that a system explain its failures in terms of the actions available to the person reading will outlast every interface technology, because it follows from who the user is rather than from how the message is rendered.

Lesson. If your users cannot act on your errors, your error handling is unfinished — and that is a property of where classification happens in the architecture, not of how the message is styled.

Every package used, in one table

Terms used in a specific sense in this package, where ordinary usage would be ambiguous.

Package What it is What it does here Considered instead
Effect key A deterministic identifier derived from (run id, step id, logical attempt), passed to a provider as an idempotency key where the action's class says it will be honoured. The mechanism that makes a retry harmless: the same key means the same operation, so a redelivery is recognised rather than repeated. "Idempotency key", which in this package means the provider-side field the effect key is carried in, not the key itself.
Replay-safety class A declared, versioned property of a connector action: idempotent (the provider honours a caller key), checkable (the effect can be read back by a natural key), or unsafe (neither). The input from which every execution guarantee is derived, and a hard release gate on publishing an action. "Retry policy", which in this package means the backoff schedule and attempt ceiling, a separate and much weaker idea.
Park A normal run state with a release time, entered on a rate limit or an unresolvable ambiguity. It consumes no worker and does not spend the retry budget. The state that lets the platform wait on somebody else's capacity without treating waiting as failure. "Failed", which is terminal and consumes the budget; a parked run has neither property.
Retry A further physical attempt at the same logical attempt, reusing the effect key. Recovery from a fault, with no new effect intended. "Replay", which is the opposite intention.
Replay An author-initiated new logical attempt, with a new effect key, re-performing the effect deliberately from the retained trigger event. Recovery of missed work, with a new effect explicitly intended and confirmed. "Retry", which must never produce a second effect.
Step ledger The append-only store of step attempts — intent before each outbound call, outcome after — keyed by run, step and attempt. The platform's system of record for run state and the point a lost worker resumes from. Run status is a projection of it. "Run history", which is the rebuildable author-facing projection, not the authority.
Trigger event log The durable partitioned record of every accepted trigger event, retained for the replay window. The platform's system of record for what happened, and the only thing a replay reads. The providers, which remain the system of record for what is true.
Connection One workspace's authorised grant into one provider account, carrying scopes, state and a reference to credential material held in custody. The unit of credential custody, of provider rate limiting, and of cursor and subscription state. "Workspace", which is the unit of tenancy, quota, billing and audit.
Quiet period The interval after which a connection that has emitted nothing is considered unexpectedly silent, derived from its own observed history. Turns an absence of data into a signal, which is the only way a dead subscription or a stuck cursor becomes visible. A health check on the poller, which reports success while delivering nothing.
Unsafe action An action for which the provider offers neither a caller-supplied idempotency key nor a reliable read-back. The class for which the platform cannot deliver at-most-once visible effect, and says so rather than guessing. "Dangerous" or "unsupported" — an unsafe action is fully supported; what is unavailable is the automatic resolution of an ambiguous outcome.
The package

Everything as it was delivered.

These files are served exactly as they were produced — the diagram pages keep their own house style because that is the artifact, not a rendering of it.