Document 18 min read

Architecture One-Pager

Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views

No-Code SaaS Automation Platform · Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views

The step attempt, not the run, is the unit of durability: a run is a resumable state machine whose entire state lives in an append-only step ledger outside the worker, and no step may be attempted without a stable effect key.

Almost everyone who works in an office has built one of these without calling it software. A form submission lands, a row appears in a spreadsheet, a message appears in a chat channel. Nobody wrote code and nobody deployed anything, and it has run unattended for two years — until the Tuesday it posts the same message three times, or the month it quietly stops because somebody in IT revoked a token and no human was told, or the Black Friday morning when 900 submissions trickle through over forty minutes because the spreadsheet API started answering 429. The product is trivial to describe and the architecture is not, because the platform owns almost none of the systems it depends on. Its triggers come from products that mostly cannot push. Its effects land in products whose APIs it cannot change, inside quotas it cannot negotiate, using credentials somebody else can revoke at any moment, on behalf of authors who cannot read a stack trace. Every interesting failure in this system belongs to somebody else — and every one of them still arrives as a complaint about us.

Separate what happened from what was done about it, and make the second one resumable. An authenticated provider delivery, or a polled read against a durable per-connection cursor, is committed to a partitioned trigger event log before the provider is acknowledged — so the accept path owes the caller durability and nothing else, and a total execution outage still loses no trigger. Admission deduplicates the event into at most one run per automation version, then schedules it under a per-workspace fair share across four queue classes so a bulk import cannot starve a two-step automation. A step runner leases the run and advances it one step at a time, writing an intent to the step ledger before each outbound call and an outcome after it; the ledger, not the worker's memory, is the run's state, so a lost worker resumes at the last committed step rather than repeating completed effects. Every outbound call leaves through a quota governor that owns the rate-limit decision, honours the provider's own backoff headers as authoritative, exchanges a short-lived single-connection token with a credential custody service in a separate account, and exits through a published NAT range customers can allowlist. A rate limit parks the run with a release time rather than failing it. An ambiguous outcome is resolved by the action's declared replay-safety class: retried under the same effect key when the provider honours one, read back first when a natural key exists, and escalated to the author when neither is true. Credential death, schema drift and runaway automations are detected, classified into an author-facing taxonomy, held rather than dropped, and surfaced in a repair inbox — because the default failure of this platform is silence, and silence has no natural end.

What it is, and what it is not

  • A resumable state machine whose state lives in an append-only ledger — not A worker that holds a run in memory and retries it from the beginning
  • At-least-once execution with at-most-once visible effect, where the provider allows it — not A blanket exactly-once promise the providers cannot support
  • A declared replay-safety class per action, authored and versioned — not One retry policy applied uniformly to forty thousand different operations
  • A rate limit as a normal run state with a release time — not A rate limit as an error that consumes a retry budget
  • Credential custody in a separate account, reached by token exchange — not A secret store the execution plane can read
  • Trigger ingestion that owns the subscription lifecycle and watches for silence — not A webhook endpoint that assumes no news is good news
  • A failure taxonomy written in the author's vocabulary with a repair path — not A provider error string forwarded to somebody who cannot read it
  • Automation of other people's SaaS products on an end user's behalf — not Orchestration of our own services, which is an adjacent use case

The decisions that are the architecture

  1. The step attempt is the unit of durability (ADR-01) — A retry, a park for a rate limit, a pause for a dead credential and a resume after a lost worker all become the same mechanism — moving a run between states in the ledger — instead of four special cases, and none of them repeats a step that already succeeded.
  2. The trigger event log is our own system of record (ADR-02) — Retention stops being a storage preference and becomes a replay commitment: the log is the only thing a replay reads, so thirty days is a promise about how far back a repair can reach, not a backup policy.
  3. Accept and execute fail independently (ADR-03) — The push endpoint acknowledges after the durable commit and before any execution, which is how a 250 ms acknowledgement fits inside every common provider delivery timeout while the run itself may take minutes — and why an execution outage loses no trigger.
  4. Every action declares its replay-safety class (ADR-04) — Idempotent, checkable or unsafe is a stored, versioned property of the action and a hard release gate, because every execution guarantee downstream is derived from it and an unclassified action would silently degrade the platform's correctness promise.
  5. A retry reuses the effect key; a replay mints a new attempt (ADR-05) — The difference between 'try again' and 'do it again' is made structural rather than procedural, so the same code path cannot accidentally deliver one when the author asked for the other.
  6. An unsafe ambiguous outcome is never retried automatically (ADR-06) — Where the provider offers no idempotency key and no read-back, the platform cannot make the step safe, so it parks and asks rather than silently choosing a duplicate invoice or a missing one on the author's behalf.
  7. No runner calls a provider directly (ADR-07) — The quota decision, the credential exchange and the egress address are one boundary, which is what keeps long-lived credentials out of the execution plane and makes the platform's rate-limit behaviour a property rather than a convention.
  8. A rate limit is a park, not a failure (ADR-08) — 429 is the ordinary case at this scale, not an incident: the run moves to a waiting state with a release time, consumes no worker, and does not spend the step's retry budget on somebody else's capacity planning.
  9. Egress is a published, stable address range (ADR-09) — Customers and providers allowlist it, which makes the range a product commitment that cannot change without notice — and makes one workspace's abusive automation everybody else's problem, which is why runaway detection is in the same decision.
  10. Push where available, poll everywhere else, and own the subscription (ADR-10) — Roughly a fifth of the catalogue can push. Subscription renewal is a component rather than a setting, because a subscription that expired unnoticed is this platform's most common invisible failure and it presents as health.
  11. A cursor advances only after durable commit (ADR-11) — Advancing first is the one-line bug that produces a permanent, silent gap. The cursor store is small, hot, and the only store here whose loss cannot be recovered from the two authoritative ones.
  12. Credential custody is a separate account (ADR-12) — The threat is not that we lose our own data; it is that a compromise of our execution plane becomes a compromise of the customers' other SaaS products. An account boundary is the blast-radius boundary.
  13. Refresh is single-flighted; credential death is terminal (ADR-13) — A thousand concurrent runs on one connection must not produce a thousand refreshes and a provider-side rotation race — a self-inflicted credential death — and a revoked token is never retried, only parked and reported.
  14. Fair scheduling, with runaway detection (ADR-14) — A bulk import is slowed, never permitted to starve another workspace, and an automation triggering itself is throttled within a minute rather than consuming a quota and a provider's goodwill.
  15. Idle must be nearly free (ADR-15) — Most of twelve million automations fire rarely. Near-zero idle cost is an architectural constraint that rules out per-automation reserved capacity, dedicated workers by default, and any design that holds a connection open per automation.
  16. Every failure carries an author-facing cause (ADR-16) — A failure class with no sentence the author can act on is an unfinished feature, not a support ticket — which makes the failure taxonomy a deliverable of the architecture rather than a property of the UI.

Why this should still be right in ten years

The cloud, the queue, the container runtime and most of the eight thousand connectors will be replaced inside this platform's life. These are the properties that should outlast them.

  • Externalised run state is older than any of this technology. Putting a long-running process's state in a durable store rather than a worker's memory is the same decision a transaction log, a saga and a workflow engine all make. The specific store will change; the property that a worker may die at any instant without repeating a completed effect will not, because it follows from the fact that the effects are in somebody else's system and cannot be rolled back.
  • The effect key outlives the protocol. Idempotency keys are an HTTP convention today and were a message-deduplication identifier before that. What endures is the requirement that an operation carry a caller-supplied identity so a redelivery can be recognised as one. Any future transport will need it for the same reason, and the platform's derivation from (run, step, logical attempt) is transport-independent.
  • Replay safety is a property of the provider, not of our code. Classifying actions by whether a retry is safe will stay necessary for exactly as long as the platform integrates systems it does not control — which is permanently. The class may be discovered automatically one day rather than declared, but something will still have to carry the answer, and the architecture's dependence on it will not change.
  • Being a good citizen of a quota is a durable constraint. Providers will always meter access, and the penalty for overshoot will always be worse than the cost of pacing. A governance point that owns the outbound decision survives every change of rate-limit algorithm, because what it encodes is that the decision belongs in one place rather than in every caller.
  • Silence will always be the hard failure. In a system whose job is to act on somebody else's behalf, the costly failure is not an error but an absence — and no amount of better tooling makes an absence announce itself. A liveness expectation per connection and a loop that closes on recovery are the answer now and will be the answer on whatever this is rebuilt on.
  • The author will not become an engineer. The product's premise is that a non-engineer can build this. That constrains the architecture permanently: every failure must be reducible to a classified cause and a repair action, which rules out designs whose internal states cannot be explained in a sentence.

Non-functional targets

Every figure below is a stated assumption. They are listed with the mechanism that is supposed to deliver them and the view where that mechanism is drawn, so a reviewer can disagree with the number and follow it to the thing that depends on it.

Quality Target How it is met View
Trigger ingest availability ≥ 99.99% monthly Stateless push and poll tiers across three AZs, durable commit before acknowledgement, independent of the execution plane 13
Push acknowledgement latency p99 ≤ 250 ms Signature check and durable log commit only; no execution, binding or credential work on the accept path 12
Execution plane availability ≥ 99.95% monthly Leased runs resumable from the ledger, so a worker or AZ loss is a resume rather than an outage 15
Start latency, push-triggered p50 ≤ 800 ms, p95 ≤ 3 s, p99 ≤ 15 s Admission reads the log continuously and leases against a per-workspace fair share across four queue classes 02
Polled trigger detection Within one interval + 30 s at p95 Jittered per-connection scheduling at 1, 5 or 15 minute entitled intervals, with change-rate backoff for quiet connections 13
Throughput 12,000 step attempts/s steady, 48,000/s burst Horizontally scaled runners; the two synchronous ledger writes per step are the binding constraint 15
Duplicate visible effects ≤ 1 per 1,000,000 attempts (idempotent and checkable classes) Deterministic effect key per attempt, read-back before retry where no key is honoured 14
Silent run loss Zero tolerated Every accepted event reaches a terminal state within 7 days or appears in the repair inbox with a named cause 18
Outbound rate-limit rate ≤ 0.1% of calls receive a 429 Per-connection and per-connector quota leases at the governor, with provider backoff headers authoritative 17
Unclassified failure share ≤ 2% of author-facing failures Typed taxonomy with a named next action per class; above the threshold the taxonomy itself is the defect 21
Durability RPO 0 for accepted events and committed ledger entries Synchronous three-AZ replication on both authoritative stores; RPO 60 s for counters and projections 10
Recovery RTO 5 min in-region; 2 h to rebuild run history Run index is a projection of the ledger and is rebuilt without a maintenance window, degraded but available 09
Tenant isolation No workspace > 5% of a shared pool for > 60 s Fair scheduling with per-workspace concurrency leases and separated queue classes 02
Credential revocation Effective within 60 s for new runs Custody marks the connection dead; tokens already issued are scoped to one run and expire with its deadline 20
Cost ≤ $0.55 per 100,000 step attempts; ≤ $0.004 per idle automation per month Serverless runners with no per-automation reservation; ledger write held to ≤ 18% of per-step cost 17

Scope

In scope

  • The automation definition as immutable versioned configuration — trigger, ordered steps, reference-based data binding, branching, per-step error policy and the connector versions it was authored against — with publish-time validation and a 60-second rollback
  • Connector manifests as versioned declarative configuration, with input and output schemas, authentication scheme, rate-limit profile and a declared replay-safety class, plus a sandbox for the exceptional case of custom code
  • Trigger ingestion in four classes — provider push, polled, scheduled and inbound catch-hook or mail — with subscription lifecycle ownership, per-connection cursors, delivery authentication, dedup and poison quarantine
  • A durable partitioned trigger event log as the platform's own system of record for what happened, retained as the replay window
  • Run admission with per-event dedup, per-workspace fair scheduling, plan quotas and separated queue classes
  • Durable step execution against an append-only step ledger, with leases, effect keys, retry with backoff and jitter, run deadlines, bounded loops and checkpointed fan-out
  • Egress governance: per-connection and per-connector quota leases, provider backoff honoured as authoritative, park-not-fail on rate limits, per-provider circuit breaking and a published egress range
  • Credential custody as a separate plane: narrowest scopes, per-workspace encryption, single-flight refresh, short-lived single-connection tokens, credential-death detection and revocation
  • Author-facing visibility: classified run history, a repair inbox grouped by cause, author-initiated replay with the duplicate consequence stated, and notification on durable failure classes
  • Multi-tenancy, plan quotas, runaway detection and per-workspace data classification driving retention, residency and redaction

Explicitly out of scope

  • Generic DAG orchestration of the organisation's own internal services, which is the distributed-workflow-orchestration-platform use case in this practice
  • Outbound webhook fan-out to subscriber endpoints, which is the webhook-delivery-service use case
  • Request-level idempotency and deduplication as a standalone service, which is the idempotency-and-dedup-service use case
  • The connectors themselves beyond the manifest contract and the sandbox they run in — building eight thousand integrations is a programme, not an architecture
  • The authoring UI's interaction design, beyond the structural requirements it imposes (reference binding, publish validation, confirmed test writes)
  • Billing, pricing and plan definition beyond the quota and fairness mechanics the architecture must enforce
  • The providers' own systems, their uptime, their schemas and their rate-limit policies, all of which are constraints rather than components

What a four-week prototype should prove

The prototype's job is to falsify the three central claims — that step-level durability really does make a worker loss invisible in somebody else's system, that the effect key genuinely prevents a duplicate against a real provider, and that parking on a 429 is cheaper than failing — on three connectors and a few thousand runs, not to build a platform.

  1. Three real connectors chosen for contrast: one that pushes and honours an idempotency key (idempotent), one that must be polled and offers a read-back by natural key (checkable), and one that neither pushes nor keys — a chat or mail action (unsafe), so all three execution contracts are exercised rather than assumed
  2. A step ledger with intent-before-call and outcome-after-call, leases, and resume-at-last-committed-step, on one shared queue with a per-workspace concurrency lease
  3. A quota governor in front of all three connectors, honouring Retry-After, with a parked-release queue separate from the interactive one
  4. Credential custody as a separate process with per-connection single-flight refresh and run-scoped token issuance — the account boundary can wait, the interface cannot
  5. A trigger path per connector: signed push for the first, a durable per-connection cursor advanced only after commit for the second
  6. A classified failure taxonomy with a repair inbox, covering only the six classes the three connectors can actually produce
  7. Cost instrumented per 100,000 step attempts with the ledger writes broken out as their own line
  • Runners are killed mid-step at a measured rate under sustained load: the acceptance criterion is zero duplicated and zero skipped effects verified against the providers own records, not against our logs — this is the falsification test for ADR-01 and ADR-05 together
  • A deliberate 429 storm against the polled provider: parking must hold the run without occupying a worker, the retry budget must be untouched, and throughput must recover at the providers stated release time rather than at a multiple of it
  • The unsafe action is driven to a genuine ambiguity by severing the connection after the request is sent: the run must park with an accurate description of what may or may not have happened, and no automatic retry may occur
  • A credential is revoked at the provider mid-run: measure the elapsed time from revocation to the author being told, and whether the parked backlog is still worth releasing when it finally is
  • The cursor store is wiped for the polled connection: measure exactly how many duplicates or how large a gap results, because that number decides how durable the cursor store has to be and is the open half of Core Architecture Question 1
  • One workspace enqueues 50,000 runs while another runs a two-step automation: the second workspaces start latency must stay inside its budget, which is the only honest test of the fairness claim
  • A thousand concurrent runs hit an expired access token on one connection simultaneously: exactly one refresh must reach the provider, and the connection must survive — the falsification test for ADR-13

Open risks, carried rather than hidden

Risk If it lands Response
The step ledger's write rate becomes the platform's ceiling Two synchronously replicated writes per step at 12,000 steps/second is the hottest path in the system; if it cannot scale, the critical design decision has to be weakened to one write or to batched commits, which reintroduces the ambiguity it was adopted to remove Measure it first in the prototype with the ledger cost broken out; shard by run id; treat any proposal to batch the intent write as a change to ADR-01 rather than an optimisation
The replay-safety class is self-declared by connector developers whose incentive is to ship A wrong declaration is indistinguishable from a correct one until an effect is duplicated in a customer's system, and the platform's correctness promise silently degrades across the catalogue Make the class a hard release gate, test it in contract tests against provider sandboxes, and treat a confirmed misclassification as a connector incident with a deprecation, not a bug fix
Retaining thirty days of trigger payloads concentrates customer data from every connected product The platform becomes the custodian of a month of its customers' data across their whole SaaS estate — a breach, residency and retention surface far larger than its own records Core Architecture Question 5 is open: hash-and-reference with re-fetch is the alternative. Per-workspace no-payload-retention is offered now, with the repair path's limits stated honestly
The shared NAT egress range is a shared reputation One workspace's abusive or runaway automation can get the range rate-limited or blocked by a provider for every other workspace behind it Runaway detection within 60 s, per-connection and per-connector quota leases, and the ability to move a workspace to a separate egress range as an incident action
Partner connector code runs beside the execution plane The largest residual security risk in the set; a sandbox escape reaches the runs, and from there the token exchange Declarative connectors as the default and code as the exception; isolated function runtime with no credentials in scope, declared egress hosts and hard resource ceilings. Core Architecture Question 7 is whether to accept partner code at all
Parked backlogs become useless or dangerous while waiting for a human A week of parked runs released at once can flood a provider and perform work nobody wants any more, turning a repair into a second incident Rate-limited release, a run deadline that terminates rather than parking indefinitely, and Core Architecture Question 6 left open on whether release is automatic or confirmed
The quota governor is on the path of every outbound call A single hot dependency whose failure stops all effects platform-wide, and whose added hop is paid by every one of 12,000 steps a second Core Architecture Question 2 is open between central arbitration and per-connection leases; the lease design degrades to local decisions with bounded overshoot if the governor is unavailable

The reasoning behind every component and technology choice is in the Architecture Decision Record: 16 records across 7 areas, each with the alternatives that lost and what the choice costs.