Architecture Decision Record
Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views
No-Code SaaS Automation Platform · Solution Architecture v1.0 · Amazon Web Services · Integration Platform Architecture · 2026-10 · 21 views
The argument these decisions serve is summarised in the Architecture One-Pager.
Sixteen decisions, grouped by the question they answer, each with the forcing question, the alternatives, what would flip the choice, and why it should outlast the technology it is realised on.
Status of this document. This is a design, not a report on a running system. Every rate, latency, ratio, threshold and retention figure is a stated assumption, chosen to be defensible and arguable rather than measured. The operating context assumed throughout is a multi-tenant automation platform with 2,500,000 active workspaces, 8,000 connectors exposing 40,000 actions and triggers, 12,000,000 enabled automations and roughly 1.8 billion step attempts a month; 4,000 trigger events a second accepted at steady state with a 4x ten-minute burst, and 12,000 step attempts a second rising to 48,000; 2,000,000 polled connections at an average five-minute interval, which is about 6,700 provider polls a second. Where the requirement supplied no number, one was invented and marked as an assumption in ask.md, section by section, so a reviewer can disagree with a figure and follow it to the decision that depends on it.
How to read a record
- Question: The forcing question: why a decision was needed at all.
- Context: The requirement, the scale and the constraint that make it hard.
- Decision: What this architecture does, stated so it can be checked.
- How it is realised on AWS: The concrete mechanism: which service or package, configured how, in which subscription.
- Options weighed: Chosen, rejected, deferred, or right elsewhere, with the reason for each.
- Consequences: What the choice buys and what it costs, both kept visible.
- Choose differently when: The conditions that would flip the decision for your system.
- Why it holds up over time: What keeps the decision right as scale, staff and technology change.
- Lesson: The principle that transfers beyond this platform.
Decision map
The durability boundary: The decision that defines the architecture, and the two that follow from it directly.
- ADR-01 · The step attempt, not the run, is the unit of durability
- ADR-02 · The trigger event log is the platform's own system of record for what happened
- ADR-03 · Acceptance and execution are separated by the log and fail independently
Effects in other people's systems: What a retry means when the thing being retried created an invoice somewhere we do not control.
- ADR-04 · Every action declares a replay-safety class, and publication is gated on it
- ADR-05 · A retry reuses the effect key; a replay mints a new logical attempt
- ADR-06 · An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically
Living inside somebody else's quota: Rate limits we did not set, cannot negotiate, and are punished for ignoring.
- ADR-07 · Every outbound call leaves through a quota governor; no runner calls a provider directly
- ADR-08 · A rate limit is a park with a release time, not a failure
- ADR-09 · Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation
Triggers from systems that cannot push: How the platform finds out that something happened, and how it notices that it has stopped finding out.
- ADR-10 · Push where available, poll everywhere else, and own the subscription lifecycle as a component
- ADR-11 · A cursor advances only after durable commit, and absence of data is itself monitored
Credentials held on somebody's behalf: Custody of millions of delegated grants into other companies' products.
- ADR-12 · Credential custody is a separate account; runners receive single-connection, single-run tokens
- ADR-13 · Refresh is single-flighted per connection, and credential death is terminal rather than retryable
Tenancy, fairness and the long tail: Twelve million automations, most of which do nothing most of the time.
- ADR-14 · Fair scheduling on shared capacity, with runaway detection as part of the same decision
- ADR-15 · Near-zero idle cost is an architectural constraint, not an optimisation
Legibility: The author is not an engineer, and that is an architectural constraint rather than a UX preference.
- ADR-16 · Every failure carries a classified, author-facing cause and a repair path
Technology by capability
Amazon Web Services was chosen for this exercise for two reasons that point the same way. The first is rotation: across this repository's use cases self-hosted open source carries 21 of 47 documented stacks and Microsoft Azure 13, with AWS on 9 — and Google Cloud, lower still at 7, took three of the seven most recent documents. Reaching for the same cloud each time teaches a service catalogue rather than architecture. The second is fit, in a weak sense that is worth stating honestly: this topic belongs to no cloud at all. Its hard parts — custody of delegated third-party credentials, a governor for quotas nobody hands you, a stable published egress identity, and a sandbox for partner code — are things every cloud leaves you to build. What AWS supplies well is the shape underneath: a very large fleet of short-lived, bursty, egress-heavy workers behind a fixed NAT range, with a partitioned durable log beside a high-write key-value store. Everything in ask.md is written vendor-neutrally — 'durable event log', not a product name — and the table below is where the neutral capability meets a specific service, with what would be used instead on another stack.
| Capability | Choice | Origin | Credible alternative | Why this one | Record |
|---|---|---|---|---|---|
| Trigger event log — the platform's system of record for what happened | Amazon MSK, partitioned by connection, 30 days retained as the replay window with tiered storage | Amazon Web Services | Confluent Cloud or self-managed Kafka; Azure Event Hubs with Capture; Google Pub/Sub with a replay subscription | Per-connection ordering without global ordering, and a retention window that is a product promise about how far back a repair can reach | ADR-02 |
| Step ledger — run state, written twice per step | DynamoDB, keyed (run_id, step_id, attempt), three-AZ synchronous replication, on-demand capacity | Amazon Web Services | Cloud Spanner or Bigtable; Azure Cosmos DB; CockroachDB or Cassandra self-hosted | The hottest write path in the platform is a keyed append with no cross-row transaction; a relational store would buy consistency the access pattern does not need at a cost it cannot absorb | ADR-01 |
| Push ingestion endpoints | ALB in front of Fargate services, per-connection signature verification, commit before acknowledgement | Amazon Web Services | Azure Front Door with Container Apps; Google Cloud Load Balancing with Cloud Run; Envoy on Kubernetes | A stateless tier whose only obligation is to authenticate and durably commit inside the provider's delivery timeout | ADR-03 |
| Trigger cursors and subscription state | DynamoDB, one small hot item per connection, advanced only after durable commit | Amazon Web Services | Azure Cosmos DB; Cloud Bigtable; Redis with AOF persistence, accepting the durability trade | Two million connections of small, very hot, very frequently updated state — and the one store whose loss is not recoverable from the two authoritative ones | ADR-11 |
| Run admission and queue classes | SQS, four queues by latency class, with per-workspace concurrency leases held in DynamoDB | Amazon Web Services | Azure Service Bus; Google Cloud Tasks; NATS JetStream or RabbitMQ self-hosted | Separated queues make the fairness guarantee structural: a bulk import cannot share a queue with a two-step automation even by accident | ADR-14 |
| Step runners | Fargate tasks, autoscaled on queue depth and age, holding no run state between steps | Amazon Web Services | Azure Container Apps; Cloud Run; Kubernetes with KEDA | Serverless containers keep idle cost near zero across twelve million mostly-dormant automations, which rules out any per-automation reservation | ADR-15 |
| Quota governor and egress | Fargate service owning per-connection and per-connector leases, behind NAT gateways with a published Elastic IP range | Amazon Web Services | Azure NAT Gateway with Container Apps; Cloud NAT with Cloud Run; Envoy with a rate-limit service on Kubernetes | The quota decision, the token exchange and the egress identity belong at one boundary, and the address range is a customer-visible commitment | ADR-07 |
| Credential custody | A separate AWS account running the custody service, with per-workspace KMS keys and envelope-encrypted material | Amazon Web Services | Azure Key Vault in a separate subscription; Cloud KMS in a separate project; HashiCorp Vault or OpenBao with a transit backend | The blast-radius boundary is an account boundary, not a service boundary: a compromised task role in the platform account must not reach key material | ADR-12 |
| Connector sandbox | AWS Lambda per invocation, no credentials in scope beyond the connection, declared egress hosts, hard CPU and memory ceilings | Amazon Web Services | Azure Functions; Cloud Functions; gVisor or Firecracker microVMs on Kubernetes; a WASM runtime for declarative-only connectors | Partner code must not share a process, a filesystem or a network namespace with the run it serves | ADR-04 |
| Definition registry and tenancy | Aurora PostgreSQL, immutable automation versions with a published-version pointer | Amazon Web Services | Azure Database for PostgreSQL; Cloud SQL or Spanner; PostgreSQL with Patroni self-hosted | Publication is strongly consistent, low-volume and relational; the hot path never reads it without a cache | ADR-05 |
| Connector catalogue | DynamoDB for manifests and version pointers, S3 for content-addressed artefacts | Amazon Web Services | Azure Cosmos DB with Blob Storage; Firestore with Cloud Storage; an OCI registry with MinIO | Hot reads of varied-shape documents, with immutable artefacts addressed by digest so a pinned version is genuinely the same bytes | ADR-04 |
| Run history and analytics projections | DynamoDB run index with a TTL by retention plan, S3 and Athena for cold history and cost analysis | Amazon Web Services | Azure Data Explorer; BigQuery; ClickHouse self-hosted | Both are rebuildable from the ledger, so their schema can change without a migration and their loss is a degradation rather than an incident | ADR-16 |
The decisions, and the alternatives that lost
The durability boundary
The decision that defines the architecture, and the two that follow from it directly.
ADR-01 · The step attempt, not the run, is the unit of durability
Status: Accepted · Shown on views: 12, 14, 07, 11
When a worker dies halfway through a five-step automation that has already created an invoice and sent an email, what happens next — and who pays for it?
Context. A run in this platform is a sequence of effects in systems the platform does not own and cannot roll back. If run state lives in the worker's memory, the only recovery from a lost worker is to re-run from the beginning, which means re-performing every step that already succeeded. For a pipeline of pure computation that is a performance problem. Here it is a wrong answer in a customer's accounting system. The same argument applies to every other interruption: a rate limit that needs waiting out, a credential that died mid-run, a provider returning 503. Each of them needs the run to stop and later continue from exactly where it was, and none of them can be served by a design whose only verb is 'start over'. At 12,000 step attempts a second the cost of externalising that state is not theoretical either, which is why this is a decision rather than an obvious default.
Decision. A run is a resumable state machine whose complete state lives in an append-only step ledger outside the worker. Every step attempt writes its intent to the ledger before the outbound call and its outcome after, carrying the run, step, attempt number, resolved inputs, effect key, outcome and the provider's response identifier where one is returned. Run status is a projection of the ledger, never an independently updated field. A worker holds only a lease; losing it releases the run for another worker to resume at the last committed step. No step may be attempted without a stable effect key (ADR-05).
How it is realised on AWS. DynamoDB holds the ledger, keyed (run_id, step_id, attempt), replicated synchronously across three availability zones, with on-demand capacity and sharding by run id. Step runners are Fargate tasks that hold no state between steps; a run lease is a short-TTL item whose expiry makes the run eligible for another runner. The intent write and the outcome write are separate items rather than an update, so the ledger is genuinely append-only and the ambiguous case — intent present, outcome absent — is directly observable rather than inferred. Run status in the author-facing run index is materialised from the ledger asynchronously and can be rebuilt from it in full.
| Option | Verdict | Reasoning |
|---|---|---|
| Step-level durability in an external append-only ledger | Chosen | A park, a retry, a pause and a resume all become the same operation on the same structure, and no completed effect is ever repeated |
| Run-level durability: the worker holds the run and retries it from the start | Rejected | Dramatically simpler and needs no ledger, but re-runs step one to retry step four — unacceptable when step one created an invoice |
| Checkpoint only at explicit author-declared boundaries | Rejected | Pushes a correctness decision onto a non-engineer who has no way to reason about it, and leaves the default unsafe |
| A single mutable run row updated in place per step | Rejected | Loses the intent-without-outcome signal that makes an ambiguous effect detectable, which is the case the whole design exists to handle |
| An embedded workflow engine holding state in its own store | Rejected | A reasonable alternative that moves the same decision behind a dependency; rejected because the effect-key semantics in ADR-05 are specific enough that owning the ledger is simpler than bending somebody else's |
What it buys
- A worker loss, a rate-limit park, a credential pause and a transient retry are one mechanism — moving a run between states — rather than four special cases with four recovery paths
- The ambiguous outcome is directly observable as an intent with no matching outcome, which is what makes ADR-06 implementable at all
- Run history, the repair inbox and author-initiated replay are all reads of one structure, so they cannot disagree with each other
- The run index can be dropped and rebuilt, which turns a schema change to author-facing history into a routine operation
What it costs
- Two synchronously replicated writes per step attempt, which at 12,000 steps/second is the platform's hottest path and its primary scaling constraint
- A priced component of every step: the ledger write is held to ≤ 18% of per-step cost, and that ceiling is a design constraint rather than a target
- Latency added to every step by the intent write before the call, paid even on the overwhelming majority of steps that succeed first time
- A lease mechanism with its own failure modes — a slow worker whose lease expires while its call is in flight is a duplicate risk that ADR-05 has to absorb
Choose differently when. Choose run-level retry instead when the effects are genuinely idempotent end to end, when runs are short enough that re-running from the start is cheap, or when the platform owns the systems the steps touch and can roll them back. A platform automating only its own internal services has all three properties, which is exactly why the adjacent distributed-workflow-orchestration-platform use case can make a different choice.
Why it holds up over time. Externalising long-running process state into a durable log is the same decision a transaction log, a saga and a workflow engine each make, and it has survived every generation of the technology underneath it. The store will be replaced; the property that a worker may die at any instant without repeating a completed effect follows from the effects being irreversible and somebody else's, which will not change.
Lesson. Decide what the unit of durability is before anything else, and derive the rest from it. If the thing you are retrying is irreversible and lives in a system you do not own, the unit cannot be the whole job.
ADR-02 · The trigger event log is the platform's own system of record for what happened
Status: Accepted · Shown on views: 09, 10, 13
When an author asks us to re-run last Tuesday's missed work, what do we read — our copy of the event, or the provider's current state?
Context. The providers are the system of record for what is true. They are not a reliable system of record for what happened: a webhook delivery is usually not retrievable afterwards, a polled change may have been superseded, and a deleted record simply is not there. If the platform keeps no copy, then a replay can only re-read the provider's current state, which means a repair performed a week after a credential died may run against data that has changed or vanished — or may be impossible because the credential is still dead, or because the provider is down. Keeping a copy makes replay exact and independent, and makes the platform the custodian of a month of its customers' data from every product they have connected.
Decision. Every accepted trigger event is durably committed to a partitioned, append-only event log before the provider is acknowledged, and retained for a declared replay window of 30 days. The log is the only thing a replay reads. Retention is therefore a recovery commitment — a promise about how far back a repair can reach — and not a storage preference.
How it is realised on AWS. Amazon MSK, partitioned by connection so per-connection ordering holds without requiring global ordering, with 30 days of retention using tiered storage so the window is affordable. Payload bodies above a size threshold are written to S3 and the log carries a reference, which keeps the log's partitions small and puts the residency and classification rules on object storage where they are easier to enforce. Admission reads the log continuously and deduplicates into at most one run per (automation version, event id).
| Option | Verdict | Reasoning |
|---|---|---|
| Retain the payload for a declared replay window; replay reads the log | Chosen | Replay is exact and independent of the provider's availability, credential state and current data |
| Retain only metadata and re-fetch the payload from the provider on replay | Rejected | Holds almost nothing and materially reduces the breach surface, but makes replay silently run against changed data and impossible when the credential is dead — which is the most common reason a replay is needed |
| Retain nothing; a missed trigger is simply missed | Rejected | Coherent and honest, and would remove a whole class of risk, but it deletes the repair journey the product is largely sold on |
| Let the workspace choose the window, including zero | Chosen | Adopted as a per-workspace data-classification setting alongside the default, with the repair path's reduced capability stated plainly rather than discovered |
What it buys
- A repair performed days later runs against exactly what arrived, not against what the provider happens to hold now
- The accept path owes the provider durability and nothing else, which is what makes a 250 ms acknowledgement achievable
- Admission, dedup and fair scheduling can all be rebuilt or re-run from the log without involving any provider
What it costs
- Thirty days of customer data from every connected product concentrated in one platform — a residency, retention and breach surface far larger than the platform's own records
- Retention is now a throughput obligation too: the log must be readable fast enough that a bulk replay does not starve live admission
- A per-workspace no-retention setting creates a second, weaker repair path that has to be explained rather than hidden
Choose differently when. Choose metadata-plus-re-fetch when the payloads are highly sensitive, when providers offer a stable read-by-id that survives the retention period, and when replays are typically attempted within minutes rather than days. Choose to retain nothing when the product makes no repair promise at all.
Why it holds up over time. Separating 'what is true' from 'what happened' is a distinction every event-driven system eventually discovers, usually after trying to reconstruct history from current state and failing. The storage will change; the need for an independent record of the event, owned by the party that must act on it, will not.
Lesson. If you promise to replay something, you must own a copy of it. A replay that reads the source is not a replay, it is a new run wearing the same name.
ADR-03 · Acceptance and execution are separated by the log and fail independently
Status: Accepted · Shown on views: 02, 12, 13
The execution plane is down. Does the platform tell the provider to go away, or take the event anyway?
Context. Provider push deliveries have short timeouts and limited, inconsistent retry behaviour; many providers retry two or three times and then drop the event permanently. A rejected delivery is therefore usually gone for good, and no amount of later recovery brings it back. Execution, by contrast, can be arbitrarily delayed without losing anything, because the work is still described in the log. The two sides of the platform have fundamentally different costs of failure, and giving them one availability target would either over-engineer execution or under-protect ingest.
Decision. The accept path authenticates the delivery, commits it durably to the event log, and acknowledges — and does nothing else. No binding, no credential resolution, no quota check and no execution happens before the acknowledgement. Ingest and execution carry separate availability targets (≥ 99.99% and ≥ 99.95%), scale independently, and a total execution outage loses no trigger.
How it is realised on AWS. ALB in front of stateless Fargate push services; the only work before acknowledgement is signature verification against the per-connection secret and the MSK produce with acknowledgement from all in-sync replicas. Admission runs as a separate consumer group entirely, so its failure, backlog or redeployment is invisible to the provider. Poll workers follow the same contract: read, commit, then advance the cursor (ADR-11).
| Option | Verdict | Reasoning |
|---|---|---|
| Authenticate, commit, acknowledge — nothing else on the accept path | Chosen | The only design in which ingest availability is genuinely independent of everything downstream |
| Execute the first step synchronously before acknowledging | Rejected | Gives the author a faster perceived start and ties the acknowledgement to credential and provider availability, which is precisely the coupling that loses events |
| Validate the payload against the automation's bindings before acknowledging | Rejected | Catches author errors a few seconds earlier at the cost of making the accept path depend on the definition registry; the same error is caught at publish time instead (ADR-16) |
| Reject deliveries when the execution backlog is beyond a threshold | Rejected | Protects the platform by discarding exactly the events it exists to not lose; back-pressure is declared through widened latency budgets instead |
What it buys
- A 250 ms p99 acknowledgement fits comfortably inside every common provider delivery timeout, while a run may legitimately take minutes or park for hours
- Execution can be redeployed, scaled, or fail entirely without a single trigger being lost
- Ingest is stateless apart from cursors, so it scales horizontally against burst without coordination
What it costs
- The author's first feedback moves from the acknowledgement to the run history, so perceived latency has to be managed explicitly as a product surface
- An automation whose definition is broken still accepts and logs events, producing runs that fail immediately and need a classified cause rather than a rejection
- Two availability targets mean two operational postures, and the ingest tier must be genuinely over-provisioned rather than merely autoscaled
Choose differently when. Collapse the two when deliveries are reliably retried by the source for hours, or when the platform owns the source. A system reading from its own durable queue has no reason to separate acceptance from execution, because nothing is lost by refusing.
Why it holds up over time. The principle that the acceptance of work and the performance of work have different costs of failure, and therefore different availability targets, outlives any particular queue. It is the same reason a payment gateway acknowledges before settling.
Lesson. Find the step where failure is irreversible — usually the one where an external party gives up — and make it depend on as little as possible.
Effects in other people's systems
What a retry means when the thing being retried created an invoice somewhere we do not control.
ADR-04 · Every action declares a replay-safety class, and publication is gated on it
Status: Accepted · Shown on views: 08, 14, 16
Can this step be retried safely? And who is supposed to know the answer — the platform, the connector developer, or the author?
Context. The platform's execution guarantees cannot be stronger than what the providers support, and support varies wildly across 40,000 actions. Some accept a caller-supplied idempotency key. Some offer a read-back by natural key that reliably shows whether the effect landed. Many — notably 'post this message' and 'send this email', among the most-used actions in the product — offer neither, and after a timeout the platform genuinely cannot know what happened. A single platform-wide retry policy must therefore be wrong for most of the catalogue: safe enough for one class, and either duplicating or silently dropping effects for another.
Decision. Every action carries a declared replay-safety class — idempotent, checkable or unsafe — as a versioned field of the connector manifest. The class is a hard release gate: an action cannot be published without it. Execution semantics are derived from the class rather than configured separately, and the author-facing behaviour of the unsafe class is stated in the product.
How it is realised on AWS. The class lives in the connector manifest in DynamoDB alongside the schemas and rate-limit profile, versioned with the connector and pinned by each automation to a compatible range. The release pipeline rejects a manifest without it, and contract tests run against provider sandboxes rather than mocks, because the failure being guarded against is the provider's behaviour, which a mock cannot reveal. The step runner reads the class at bind time and selects the ambiguity-resolution path in ADR-06.
| Option | Verdict | Reasoning |
|---|---|---|
| A declared, versioned per-action class, gated at release | Chosen | Puts the answer where the knowledge is, and makes it reviewable, testable and pinnable |
| One platform-wide retry policy | Rejected | Simple and wrong for most of the catalogue: necessarily either duplicating or dropping, depending on which way it is tuned |
| Infer the class at runtime from provider responses | Rejected | Attractive and unsound: the inference is only testable by causing the ambiguity it is meant to resolve, in a customer's system |
| Let the author choose per step | Rejected | Asks a non-engineer a question they have no basis to answer, and the wrong answer is invisible until an invoice is duplicated |
| Let the author choose only the unsafe class's behaviour | Chosen | Adopted in ADR-06: the author decides what to do about an ambiguity, never whether the ambiguity exists |
What it buys
- The platform can state a different, honest guarantee per action instead of one promise that is false for most of them
- The interface catalogue becomes organised by the property that actually matters, which makes the unsafe surface countable and reviewable
- Pinning a connector version pins the guarantee, so a provider change cannot silently alter execution semantics for a published automation
What it costs
- The class is self-declared by connector developers whose incentive is to ship, and a wrong declaration is indistinguishable from a right one until an effect is duplicated
- A per-action field across 40,000 actions is a large surface to review and to keep honest as providers change
- The product must explain a distinction users did not ask for, at the moment they are least interested in it
Choose differently when. Drop the classification when every integration target is under your control and can be made idempotent by fiat. An internal platform can mandate idempotency keys across its own services and needs none of this.
Why it holds up over time. For as long as a platform integrates systems it does not control, something must carry the answer to 'is a retry safe here'. The mechanism may one day be discovered rather than declared, but the dependency on the answer is permanent.
Lesson. When a guarantee depends on a third party's capability, make the capability an explicit, versioned, testable field — not an assumption spread through the code that uses it.
ADR-05 · A retry reuses the effect key; a replay mints a new logical attempt
Status: Accepted · Shown on views: 11, 12, 14
The author presses a button marked 'run it again'. Did they mean 'the first one may not have worked' or 'do it a second time on purpose'?
Context. Those two intentions produce opposite correct behaviours, and a platform that conflates them will sometimes duplicate an invoice and sometimes fail to send one. The distinction cannot live in a flag read at call time, because by then the two paths share the same code and the same provider call; it has to be built into how the key is derived, so that the difference is structural and cannot be got wrong by a caller.
Decision. Each step attempt carries a deterministic effect key derived from (run id, step id, logical attempt). A retry is a new physical attempt at the same logical attempt and therefore reuses the key, which is passed to the provider as an idempotency key wherever the action's class says it will be honoured. A replay is an author-initiated action that mints a new logical attempt and therefore a new key, is recorded in the ledger as a replay, and requires the duplicate consequence to be stated and confirmed first.
How it is realised on AWS. The key is a hash of the run id, step id and logical attempt number, computed by the runner before the intent write so that the ledger entry and the outbound call carry the same value. The governor forwards it in whatever header the connector manifest declares. The ledger's attempt key distinguishes physical attempts, while the logical attempt is a field on the entry — so 'how many times did we try' and 'how many times did the author ask for this' are separately countable.
| Option | Verdict | Reasoning |
|---|---|---|
| Key derived from (run, step, logical attempt); replay increments the logical attempt | Chosen | Makes the distinction structural: the same code cannot deliver a replay when a retry was meant |
| A random key per physical attempt | Rejected | Defeats the entire purpose; every retry becomes a new effect at the provider |
| A key derived from the payload's content hash | Rejected | Collides across two legitimately identical actions — two identical messages the author genuinely wanted twice — and silently drops the second |
| One key per run | Rejected | Cannot distinguish steps, so a provider honouring it would reject the second step of the same run |
| Let the connector supply its own key derivation | Rejected | Flexible, and makes the platform's correctness depend on 8,000 independent implementations of the same subtle rule |
What it buys
- A retry after a timeout, a lease expiry or a park cannot produce a second effect where the provider honours keys
- 'Try again' and 'do it again' are visibly different in the ledger, so the audit answers which one happened
- The key is transport-independent and survives a change of provider API or protocol
What it costs
- The key must be minted before the intent write, so the derivation sits on the hottest path and cannot depend on anything slow
- It is only useful where the provider honours it — the unsafe class carries the key and gets nothing for it, which has to be explained rather than assumed
- A lease expiry during an in-flight call produces two physical attempts at the same logical attempt, which is safe only because the key is reused — making lease handling a correctness concern, not just a scheduling one
Choose differently when. A single key per operation is enough when the system performs one effect per job and never needs an intentional repeat. The two-level scheme earns its complexity only where a deliberate re-do is a product feature.
Why it holds up over time. Caller-supplied operation identity is an old idea with new names each decade — message dedup id, idempotency key, request id. What endures is that the caller, not the callee, has to assert what counts as the same operation.
Lesson. When two user intentions demand opposite behaviours, encode the difference in the data structure rather than in a parameter, so no code path can confuse them.
ADR-06 · An ambiguous outcome on an unsafe action is parked and escalated, never retried automatically
Status: Accepted · Shown on views: 14, 21, 12
The request was sent, the connection dropped, and the provider offers no key and no read-back. Do we send it again or not?
Context. This is the one case where the platform genuinely cannot know what happened and cannot find out. Retrying risks a second email to a customer or a second message in a channel. Not retrying risks the thing never happening at all, silently. Both are wrong, and which is less wrong depends entirely on the action and the author's situation — a duplicate invoice is a serious problem, a duplicate internal reminder is not. The platform has no basis to choose, and choosing silently means being wrong without telling anybody.
Decision. For an action classified unsafe, an ambiguous outcome is never retried automatically. The run is parked in a state that names the ambiguity, the author is told exactly what may or may not have happened, and two explicit choices are offered: continue without retrying, accepting a possible gap, or retry and accept a possible duplicate. For checkable actions the platform reads back by natural key before retrying. For idempotent actions it simply retries under the same key.
How it is realised on AWS. The runner detects the ambiguity directly from the ledger's shape — an intent with no matching outcome — and branches on the manifest's class. The parked state carries a typed reason and the provider's last response if any, and surfaces in the repair inbox as its own category rather than mixed with ordinary failures. The run deadline still applies, so a parked-ambiguous run terminates into a named terminal state rather than waiting indefinitely for an author who will never look.
| Option | Verdict | Reasoning |
|---|---|---|
| Park and escalate to the author with two explicit choices | Chosen | The platform refuses to guess on the author's behalf about an effect in their own business |
| Always retry | Rejected | Produces duplicate invoices and duplicate customer emails, which is the failure customers escalate hardest and trust least |
| Never retry | Rejected | Produces a silent gap, which is worse because nobody finds out — the exact failure mode journey 05 is about |
| Let the connector developer pick the default | Rejected | Moves the guess one step away without improving the information available to make it |
| Refuse to publish automations whose first write is unsafe without acknowledgement | Deferred | Attractive for high-stakes actions and left to Core Architecture Question 3; it raises a friction cost at authoring time that has not been measured |
What it buys
- The platform never silently creates a duplicate effect in a customer's finance or communications system
- The ambiguity is named and visible rather than resolved by a default nobody chose
- The decision is made by the only party with the business context to make it
What it costs
- A parked run waits for a human, and many authors will not look — the run deadline then terminates it, which is itself a form of the silent gap this decision tried to avoid
- The product must explain a genuinely hard concept at a bad moment, and 'we do not know if your email sent' is a difficult sentence
- Parked-ambiguous runs accumulate and need their own inventory, alerting and expiry policy
Choose differently when. Pick a default and apply it when the actions are low-stakes and homogeneous — an internal notifier can always retry, and nobody will mind. The escalation is worth its friction only where the effects are irreversible and matter commercially.
Why it holds up over time. Distributed systems will not stop producing ambiguous outcomes, and no protocol removes the case where the acknowledgement is lost and the operation cannot be queried. The lasting decision is whose judgement resolves it.
Lesson. When a system cannot know the answer and the consequences of guessing fall on someone else, surface the choice rather than picking a default and calling it a policy.
Living inside somebody else's quota
Rate limits we did not set, cannot negotiate, and are punished for ignoring.
ADR-07 · Every outbound call leaves through a quota governor; no runner calls a provider directly
Status: Accepted · Shown on views: 07, 12, 19
Twelve thousand runners want to call the same provider at once. Who decides whether they may?
Context. Providers meter access per end-user account, per connector application credential, per region and per source address, and the penalties for overshoot escalate from a 429 to a suspended application credential that breaks every workspace at once. A decision made independently by each runner cannot respect a limit that is shared across runners, and the shared cases are the dangerous ones: a per-application limit is one scarce resource that every tenant is drawing on simultaneously. The same boundary is also the natural place to exchange a short-lived credential, because it is the last point before the call leaves the platform.
Decision. Every outbound third-party call passes a governance point that owns the quota decision; no step worker may call a provider on its own judgement. The governor enforces limits at each level the provider actually imposes them, honours provider backoff headers as authoritative over its own schedule, exchanges the short-lived single-connection token with custody, opens per-provider circuit breakers, and exits through the published egress range.
How it is realised on AWS. A Fargate service holding per-connection and per-connector leases in DynamoDB, in front of NAT gateways with a published Elastic IP range. Leases are checked out in blocks sized to the connection's observed rate so the common case is a local decision, with central arbitration reserved for per-application limits where tenants genuinely contend. Circuit breaker state is per provider and shared across the fleet, so one runner discovering an outage protects the rest.
| Option | Verdict | Reasoning |
|---|---|---|
| A governance point owning the decision, with leases checked out in blocks | Chosen | Correct for shared limits, and degrades to local decisions with bounded overshoot when the governor is slow |
| Per-worker local budgets with periodic reconciliation | Rejected | Fastest and overshoots exactly when traffic is burstiest, which is when a provider is least forgiving |
| A central broker consulted synchronously on every call | Rejected | Correct and puts a hard synchronous dependency plus a network hop on the hottest path in the platform |
| Rely on the provider's 429 and back off reactively | Rejected | Makes the provider's rate limiter our rate limiter, which works until the provider responds by suspending the application credential |
What it buys
- A limit shared across tenants can actually be respected and fairly apportioned, which no local scheme achieves
- Credential exchange, egress identity and the quota decision are one boundary, which is what keeps long-lived credentials out of the execution plane
- Provider health becomes a first-class platform signal rather than something each runner rediscovers
What it costs
- The governor is on the path of every outbound call and is the platform's hottest single dependency
- Leases add a tuning problem: too small and they are a central broker with extra steps, too large and they overshoot during a burst
- A cross-account token exchange per run adds latency, mitigated by caching for the run's duration at the cost of widening the revocation window
Choose differently when. Local budgets are sufficient when every limit is per end-user account, traffic is smooth, and the penalty for a modest overshoot is only a 429. Central arbitration earns its cost when tenants share one scarce credential.
Why it holds up over time. Providers will always meter, and the penalty for overshoot will always exceed the cost of pacing. Putting that decision in one place rather than in every caller survives any change of rate-limiting algorithm.
Lesson. A constraint that is shared cannot be enforced by parties acting independently. Put the decision where the constraint is, even when that means a hop on the hot path.
ADR-08 · A rate limit is a park with a release time, not a failure
Status: Accepted · Shown on views: 12, 21, 17
A provider says 'not now, try in ninety seconds'. Is that an error?
Context. At this scale a 429 is the ordinary case, not an incident — the platform makes billions of calls a month against quotas it did not set. Treating it as a failure has three bad consequences: it consumes the step's retry budget on something that is not a fault, it occupies a worker for the duration of a backoff, and it eventually exhausts the attempts and fails a run that was never actually broken. It also teaches the platform to treat the provider's own backoff signal as advice rather than instruction, which is how an application credential gets suspended.
Decision. A rate-limit response moves the run to a waiting state with a release time derived from the provider's own backoff header, which is authoritative in preference to the platform's schedule. A park consumes no worker and does not spend the step's retry budget. The connection's pacing adapts in response, so repeated parking feeds back into the lease size rather than simply repeating.
How it is realised on AWS. The governor returns a typed park outcome with a release timestamp; the runner writes it to the ledger and releases its lease. A dedicated parked-release queue class holds the run until its time, so parked work is isolated from interactive runs and cannot distort their queue-age signal. The platform's own SLO is that ≤ 0.1% of outbound calls receive a 429 at all — the governor's job is to make parking rare, not to make it cheap.
| Option | Verdict | Reasoning |
|---|---|---|
| Park with a provider-derived release time, outside the retry budget | Chosen | Treats the ordinary case as ordinary, and keeps the retry budget for actual faults |
| Treat 429 as a retryable error with exponential backoff | Rejected | Spends the retry budget on capacity rather than faults, and holds a worker through every backoff |
| Fail the run and let the author retry | Rejected | Converts a routine, self-resolving condition into manual work for someone who cannot influence it |
| Use our own backoff schedule and ignore Retry-After | Rejected | Substitutes our guess for the provider's statement about its own capacity, which is both rude and usually wrong |
What it buys
- Throughput recovers at the provider's stated time rather than at an arbitrary multiple of it
- The retry budget stays meaningful as a signal about faults, which keeps the failure taxonomy honest
- Parked runs occupy no compute, so a provider-wide slowdown costs storage and patience rather than capacity
What it costs
- A parked-release queue class with its own scheduling, inventory and alerting
- Parked runs can accumulate into a backlog that fires in a burst at release, which needs its own rate-limited release path
- The author experiences a slow automation with no error, which must be explained in the run history as 'running slowly' rather than left blank
Choose differently when. Treat rate limits as errors when they are genuinely exceptional and indicate misconfiguration rather than capacity — a well-provisioned internal API where a 429 means somebody has a bug.
Why it holds up over time. Backpressure from a dependency is information, not failure. Systems that learn to wait on it outlive systems that learn to retry through it, and that has been true of every protocol that has ever had a busy signal.
Lesson. Distinguish 'this is broken' from 'not right now'. Budgets, alerts and retries that conflate them will spend themselves on the wrong thing.
ADR-09 · Egress is a published, stable address range, and one tenant's behaviour is everyone's reputation
Status: Accepted · Shown on views: 15, 19, 07
A customer's security team wants to allowlist us. What address do we give them, and what happens when a neighbour abuses it?
Context. Enterprise customers and many providers restrict inbound API access by source address, so a platform that cannot name its egress addresses cannot be adopted by them at all. Publishing a range makes it a commitment: it cannot change without notice, which constrains every future networking decision. It also makes the range a shared identity — providers rate-limit, throttle and block by source address, so a single runaway automation can degrade or sever access for every workspace behind it.
Decision. The platform egresses through NAT gateways with a published, stable address range that customers and providers may allowlist, and treats changes to it as a product change requiring notice. Because the range is shared, runaway detection, per-connection and per-connector quota leases, and the ability to move a workspace onto a separate egress range are part of the same decision rather than separate features.
How it is realised on AWS. NAT gateways with Elastic IPs per availability zone, published and versioned as documentation. Every outbound request carries a connector-identifying user agent and a correlation identifier a provider can quote in a support ticket. Residency-pinned workspaces egress from their own region's range. Runaway detection throttles an automation within 60 s of exceeding 10x its 7-day baseline, before a provider notices rather than after.
| Option | Verdict | Reasoning |
|---|---|---|
| A published stable NAT range, with runaway controls as part of the same decision | Chosen | Enterprise adoption requires it, and the shared-reputation risk it creates has to be mitigated in the same breath |
| Ephemeral egress addresses | Rejected | Operationally simplest and excludes every customer whose security policy requires an allowlist |
| A dedicated egress address per workspace | Rejected | Removes the shared-reputation risk entirely and does not scale to 2.5 M workspaces at any sane cost |
| Dedicated egress for workspaces that require or earn it | Chosen | Adopted as both a plan feature and an incident response, on the same code path |
What it buys
- Enterprise customers can allowlist the platform, which is a precondition for a large part of the market
- A provider investigating traffic can identify the caller and quote a correlation id back, turning an opaque block into a conversation
- Residency pinning extends naturally, because the egress identity is already per-region
What it costs
- The range becomes immovable without a notice period, constraining future networking changes
- One workspace's abuse is every workspace's problem, which raises the stakes on runaway detection considerably
- Per-workspace egress as an escape hatch needs to exist before it is needed, not after
Choose differently when. Skip the published range when no customer requires allowlisting and providers do not restrict by address — a consumer-only product can egress from anywhere and avoid the shared-reputation problem entirely.
Why it holds up over time. Network identity as a trust signal has survived every change of protocol and will outlast the current generation of address-based controls, because the underlying need — naming who is calling — does not go away.
Lesson. A shared identity is a shared liability. If you publish one, you must also own the controls that stop one tenant spending the reputation of all of them.
Triggers from systems that cannot push
How the platform finds out that something happened, and how it notices that it has stopped finding out.
ADR-10 · Push where available, poll everywhere else, and own the subscription lifecycle as a component
Status: Accepted · Shown on views: 13, 08, 17
Four fifths of the catalogue will not tell us when something changes. How does the platform find out, and how does it notice that it has stopped finding out?
Context. Provider push is cheap, fast and available for roughly a fifth of the catalogue, and it brings subscription lifecycle, signature verification, duplicate deliveries and — most damagingly — silent subscription death, where a subscription expires or is revoked and the platform simply stops receiving events with no error anywhere. Polling works everywhere, costs linearly in connections, and at 2,000,000 connections on an average five-minute interval amounts to roughly 6,700 provider polls a second that somebody has to pay for and that providers have to absorb.
Decision. Both, treated as architecturally different rather than hidden behind one abstraction. Push is preferred wherever the provider supports it, and the platform owns the subscription lifecycle as a component: create, renew before expiry, detect silent death, re-create. Polling adapts its interval per connection to the provider's quota profile, the plan's entitlement and the observed change rate, with declared jitter, and the initial poll establishes a watermark rather than emitting an account's entire history.
How it is realised on AWS. A subscription manager tracks every push subscription's renewal deadline and renews ahead of it; a liveness expectation per connection raises a quiet-period alarm when a source has emitted nothing for longer than its declared quiet period, which is wired off the event log so it watches for the absence of data rather than the health of the poller. Poll workers are Fargate tasks reading cursors from DynamoDB with per-connection jitter, backing off for quiet connections and batching provider queries where the API allows.
| Option | Verdict | Reasoning |
|---|---|---|
| Hybrid, with the subscription lifecycle owned explicitly | Chosen | The only design that covers the catalogue while giving the fifth that can push the latency it deserves |
| Poll everything, for uniformity | Rejected | One mechanism and one failure taxonomy, at the cost of latency on the connectors where latency is most visible and a much larger polling bill |
| Push only, and omit connectors that cannot push | Rejected | Would cut the catalogue — the growth ceiling — by roughly four fifths |
| Treat polling as a fallback inside one trigger abstraction | Rejected | Hides two genuinely different failure taxonomies behind one interface, which is how a dead subscription gets mistaken for a quiet account |
What it buys
- Push connectors get seconds of latency and polled connectors get a declared, plan-appropriate interval, instead of both getting the worse of the two
- A subscription that expires is detected and re-created rather than becoming an invisible outage
- Jitter prevents the platform from aligning a million connections on the minute boundary against one provider
What it costs
- Two freshness models, two dedup mechanisms and two failure taxonomies behind one product promise
- Polling infrastructure is assumed at ≤ 25% of total platform compute cost, attributed to the connectors that lack push support
- Subscription lifecycle is ongoing work that scales with the catalogue and with provider API churn
Choose differently when. Poll everything when the catalogue is small and latency expectations are loose; push only when every integration target is modern and under contract. The hybrid is the price of a long tail.
Why it holds up over time. The split between systems that notify and systems that must be asked has outlived several generations of integration technology and shows no sign of closing, because it reflects the providers' priorities rather than a technical limitation.
Lesson. Do not unify two mechanisms whose failure modes differ. One abstraction over both hides exactly the distinction operators need when something stops working.
ADR-11 · A cursor advances only after durable commit, and absence of data is itself monitored
Status: Accepted · Shown on views: 13, 10, 17
What is the smallest piece of state in this platform, and what happens when we lose it?
Context. A polled connection's cursor — a watermark, an updated-at bound or a page token — is a few bytes per connection. It is also the only state in the architecture whose loss cannot be recovered from the two authoritative stores: the event log knows what arrived, and the ledger knows what was done, but neither knows what the provider would have returned. Lose a cursor forward and there is a permanent silent gap; lose it backward and there is a flood of duplicates. Worse, the ordinary failure here produces no error at all — a stuck cursor and a healthy quiet account look identical.
Decision. A cursor advances only after the events it covers are durably committed to the event log, never before. The cursor and subscription store is treated as a distinct data zone with its own durability posture rather than grouped with other operational state. Every connection carries a liveness expectation, and an unexpectedly quiet source raises an alarm — the platform monitors the absence of data, not only the health of the components that fetch it.
How it is realised on AWS. DynamoDB, one small item per connection, written after the MSK produce is acknowledged by all in-sync replicas. The quiet-period expectation is derived from the connection's own observed history rather than a global constant, so a genuinely low-traffic connection does not alarm. The alarm is evaluated off the event log, so it fires whether the cause is a dead subscription, a stuck cursor, a paused automation or a provider that has quietly stopped returning results.
| Option | Verdict | Reasoning |
|---|---|---|
| Advance after commit, with a per-connection liveness expectation | Chosen | Guarantees at-least-once at the cost of occasional duplicates, which ADR-05's effect key absorbs |
| Advance before commit | Rejected | A one-line difference that produces a permanent, silent, unrecoverable gap — the worst failure mode available to this platform |
| Make the cursor store as durable as the event log | Deferred | Removes the residual risk and roughly doubles the cost of the hottest small store; left open as the second half of Core Architecture Question 1 |
| Reconstruct cursors from the event log after a loss | Rejected | Recovers the last event seen but not the provider's pagination position, so it narrows the gap without closing it |
What it buys
- A crash between read and commit costs a duplicate, which the effect key makes harmless, rather than a gap, which nothing makes harmless
- A stuck cursor, a dead subscription and a silently broken provider all surface through one alarm that watches for silence
- The riskiest store in the architecture is named as such, which is the precondition for anyone treating it carefully
What it costs
- Duplicates are now routine on the poll path and the dedup window has to be sized for them
- A per-connection liveness baseline is state about state, and it has to be learned rather than configured across two million connections
- The residual risk is real and unresolved: a cursor store failure still produces duplicates or a gap, and the choice between them is not yet made
Choose differently when. Advancing before commit is defensible only when the source can be fully re-read cheaply and idempotently at any time, which makes the cursor an optimisation rather than state.
Why it holds up over time. Checkpoint-after-commit is the same rule as a consumer offset in a log, a replication position, or a file-read watermark. It will outlive this platform because it follows from the ordering of durability, not from any technology.
Lesson. Find the smallest piece of state whose loss cannot be reconstructed from anything else, and design for that one first. It is rarely the biggest store, and it is usually the one nobody is watching.
Credentials held on somebody's behalf
Custody of millions of delegated grants into other companies' products.
ADR-12 · Credential custody is a separate account; runners receive single-connection, single-run tokens
Status: Accepted · Shown on views: 19, 20, 07
Our execution plane is compromised. How much of our customers' other SaaS estate goes with it?
Context. The platform holds delegated grants into millions of customers' CRMs, mailboxes, file stores and finance systems. Its own data is not the prize; the access is. The execution plane is simultaneously the largest attack surface in the system — it binds untrusted provider responses, runs partner connector code, and makes outbound calls to arbitrary hosts — and the component that needs credentials most often. Those two facts pull in opposite directions, and resolving them by convenience means a single compromised task role reads every refresh token the platform holds.
Decision. Credential custody runs in a separate cloud account with its own key material, reachable only through a narrow exchange interface. No long-lived credential is ever available to a step worker. A worker receives a short-lived token scoped to one connection and one run, exchanged at the egress boundary (ADR-07); the refresh token never leaves custody. Material is encrypted under a per-workspace key, and plaintext never appears in logs, run history, error messages or the ledger.
How it is realised on AWS. A separate AWS account running the custody service on Fargate, with per-workspace KMS keys and envelope-encrypted ciphertext. Cross-account access is by a narrowly scoped role permitting only the exchange operation, never a read. The connection's state and scopes stay in the platform account so the control plane can query them; only the ciphertext and key reference live in custody. Every exchange is written to an audit record retained seven years independently of the workspace's retention setting.
| Option | Verdict | Reasoning |
|---|---|---|
| A separate account, with exchange-only access and run-scoped tokens | Chosen | The blast-radius boundary matches the account boundary, which is the only boundary an attacker cannot talk their way across with a stolen role |
| A managed secret store in the same account, read by runners | Rejected | Far simpler and makes a compromised runner equivalent to a compromise of every customer's connected products |
| A separate service in the same account | Rejected | Better than nothing and still inside one IAM blast radius, which is the thing being defended against |
| Hold no credentials: ask the author to re-authorise each run | Rejected | Eliminates the risk and the product, which exists precisely to run unattended |
What it buys
- A compromised execution plane yields run-scoped tokens for connections currently executing, not the estate
- Per-workspace keys make crypto-deletion of one workspace's credentials a key operation rather than a data sweep
- The audit of credential use is kept outside the workspace's own control, so it cannot be shortened by the party it holds accountable
What it costs
- A cross-account hop on the path of every step that touches a provider, mitigated by per-run caching at the cost of a wider revocation window
- Two accounts to operate, deploy and keep in step, with their own failure and permission-drift modes
- Custody becomes a 99.99% dependency — every step needs it — which is a higher bar than the execution plane it serves
Choose differently when. A same-account secret store is proportionate when credentials reach only systems you already own, so a compromise gains the attacker nothing they did not already have.
Why it holds up over time. Separating custody of a credential from use of a credential is the oldest idea in access control and is independent of the technology implementing either. The boundary will move between accounts, enclaves or HSMs; the separation will not go away.
Lesson. Size the blast radius by what an attacker gains, not by what you store. If the value is access to other people's systems, the boundary has to be one your own compromised code cannot cross.
ADR-13 · Refresh is single-flighted per connection, and credential death is terminal rather than retryable
Status: Accepted · Shown on views: 20, 21, 18
A thousand runs on one connection hit an expired access token in the same second. What happens?
Context. Without coordination the answer is a thousand simultaneous refresh attempts against the provider. Many providers rotate the refresh token on use and invalidate the previous one, so concurrent refreshes race: one succeeds, the rest present a token that has just been invalidated, and the connection dies — a credential death entirely manufactured by the platform. The converse error is equally damaging: treating a genuine revocation as a transient fault and retrying produces a sustained authentication storm against a provider that has already said no, which is a reliable way to have an application credential suspended.
Decision. Access-token refresh is serialised per connection with single-flight semantics: one refresh proceeds, concurrent callers wait for its result. Credential death — revocation, expiry, password change, scope removal — is classified as non-retryable. Affected runs are parked, the automation is paused after a declared grace period, and the connection's owner is notified. The platform never retries an authentication failure as though it were transient.
How it is realised on AWS. Custody holds a per-connection lock for the duration of a refresh and returns the new token to all waiters. Failure classification distinguishes an expired access token (refresh and continue) from a dead grant (terminal), using the provider's error code as declared in the connector manifest rather than inferring from a status code alone. The terminal state feeds the automation health loop: classify, hold, notify, repair, observe.
| Option | Verdict | Reasoning |
|---|---|---|
| Single-flight refresh per connection; credential death terminal | Chosen | Prevents the platform from causing the failure it is trying to recover from |
| Let each run refresh independently | Rejected | Causes refresh-token rotation races and converts a routine expiry into a dead connection under load |
| Refresh proactively on a schedule | Chosen | Adopted alongside, as an optimisation that reduces how often the single-flight path is contended; it does not remove the need for it |
| Retry authentication failures with backoff | Rejected | Produces an authentication storm against a provider that has already declined, and risks suspension of the application credential |
What it buys
- A connection under heavy concurrent load refreshes once, not a thousand times
- A revoked credential fails fast and visibly instead of degrading slowly behind retries
- Authentication traffic to providers stays proportional to connections rather than to runs
What it costs
- A per-connection lock in custody, on the path of every step, with its own contention and timeout behaviour
- Waiters block on a refresh they did not initiate, adding tail latency at exactly the moment a connection is busiest
- Distinguishing 'expired' from 'revoked' depends on provider error semantics declared in the manifest, which is another field that can be wrong
Choose differently when. Independent refresh is harmless where providers do not rotate refresh tokens and tolerate concurrent refreshes. The single-flight requirement comes entirely from rotation semantics.
Why it holds up over time. Collapsing concurrent identical work into one operation is a pattern older than the web and applies wherever a shared resource is refreshed under load. Token rotation makes it a correctness requirement rather than an efficiency one, and rotation is becoming more common, not less.
Lesson. When a recovery action has side effects on shared state, concurrency control is part of correctness. An unsynchronised refresh is not a slow path, it is a bug.
Tenancy, fairness and the long tail
Twelve million automations, most of which do nothing most of the time.
ADR-14 · Fair scheduling on shared capacity, with runaway detection as part of the same decision
Status: Accepted · Shown on views: 02, 17, 15
One workspace imports fifty thousand rows. What happens to everybody else's two-step automation?
Context. The workload is extremely uneven: most workspaces run a handful of automations occasionally, and a few run bulk operations that enqueue tens of thousands of runs in a burst. On one shared pool with a single queue, the bulk workspace wins simply by arriving, and every other tenant's latency collapses. Worse, some of those bursts are not legitimate work at all but automations triggering themselves — an automation that writes to the sheet it is watching — which can consume a workspace's quota and a provider's goodwill in minutes.
Decision. Shared capacity with fair scheduling across workspaces and per-workspace concurrency leases, so no workspace holds more than 5% of a shared pool for longer than 60 s while others have queued work. Queue classes are separated by expected latency — interactive, retry, parked-release and bulk — so a backfill never shares a queue with a short automation. Runaway detection throttles an automation within 60 s of exceeding 10x its own 7-day baseline, then pauses it with notice.
How it is realised on AWS. Four SQS queues by latency class, with per-workspace concurrency leases in DynamoDB enforced at admission rather than at the runner, so an over-quota workspace never occupies a worker it will not be allowed to use. Runaway baselines are per automation and learned from its own history, because a rate that is pathological for one automation is normal for another. Dedicated capacity for workspaces that require it runs on the same code path, as a lease pool rather than a separate fleet.
| Option | Verdict | Reasoning |
|---|---|---|
| Shared pool, fair scheduling, separated queue classes, runaway detection | Chosen | Best utilisation with a bounded tail, and keeps idle cost near zero (ADR-15) |
| One shared pool, first-come-first-served | Rejected | Simplest and gives the worst tail exactly when it is most visible |
| Pools per plan tier | Rejected | Simple isolation that strands capacity and still does not protect tenants from each other within a tier |
| Dedicated workers per workspace | Rejected | Ends noisy neighbours and breaks the near-zero idle cost that makes twelve million automations affordable |
| Dedicated capacity as a plan feature on the same code path | Chosen | Adopted for the workspaces that require it, without creating a second execution path to maintain |
What it buys
- A bulk import is slowed rather than permitted to starve, and the guarantee is structural because the queues are physically separate
- A self-triggering automation is caught in a minute, before it spends a provider's goodwill or the workspace's quota
- Utilisation stays high, which is what keeps per-step cost inside its target
What it costs
- Fair scheduling is real machinery with its own tuning, and its failure mode is a latency regression that is hard to attribute
- Per-automation baselines are state that must be learned across twelve million automations and will misfire on genuinely spiky but legitimate work
- If plan-differentiated latency is ever sold, fairness and the commercial promise are in direct conflict and the architecture has to encode which wins
Choose differently when. First-come-first-served is adequate when the workload is homogeneous and no tenant can enqueue orders of magnitude more than another. Dedicated pools win when tenant isolation is a compliance requirement rather than a performance one.
Why it holds up over time. Fair queueing across tenants on shared infrastructure is as old as time-sharing and has survived every change in what the shared resource is. The specific scheduler will change; the need to stop one tenant monopolising a shared pool will not.
Lesson. Make isolation structural rather than behavioural. Separate queues enforce a guarantee that a shared queue with good intentions cannot.
ADR-15 · Near-zero idle cost is an architectural constraint, not an optimisation
Status: Accepted · Shown on views: 15, 17, 07
Twelve million automations exist. Most of them will not run today. What do they cost?
Context. The business is the long tail: the median workspace has a few automations that fire occasionally, and a large fraction of the twelve million are effectively dormant. If a dormant automation costs anything meaningful — a reserved worker, an open connection, a provisioned queue, a polling slot it does not need — the economics fail at a scale the product must reach to be viable. This is a genuine architectural constraint rather than a cost optimisation, because it rules out whole classes of design at the point where they are chosen.
Decision. An enabled but dormant automation holds no reserved compute, no open connection and no worker slot, at an assumed cost of ≤ $0.004 per automation per month. This rules out per-automation reserved capacity, dedicated workers by default, per-automation long-lived connections and any design where an automation's existence rather than its activity drives cost. Poll cost must be sub-linear in connection count, achieved through interval adaptation, change-rate backoff for quiet connections and batched provider queries.
How it is realised on AWS. Fargate tasks autoscaled on queue depth and age, with nothing allocated per automation. Automations exist as rows in the definition registry and subscriptions or cursors in a small key-value store — bytes, not capacity. Poll scheduling adapts per connection, so a connection that has not changed in a month is polled at its plan's floor rather than its ceiling.
| Option | Verdict | Reasoning |
|---|---|---|
| Serverless workers, no per-automation allocation, adaptive polling | Chosen | The only shape where twelve million mostly-dormant automations are economically viable |
| A reserved worker or process per active automation | Rejected | Simplest mental model and fails on cost by orders of magnitude at this scale |
| Long-running connections per automation | Rejected | Would reduce latency and makes dormant automations expensive in exactly the way the economics cannot absorb |
| Archive dormant automations to cold storage | Deferred | Worth doing if registry cost becomes material; it is not, because a dormant automation is already only bytes |
What it buys
- The long tail is viable, which is the business rather than a technical nicety
- Capacity tracks actual work, so a quiet weekend costs nothing and a Monday burst scales up
- No per-automation resource means no per-automation leak, exhaustion or cleanup path
What it costs
- Cold-start latency on scale-up, paid by whichever runs arrive first after a quiet period
- A per-task premium over long-running instances, accepted deliberately as the price of the constraint
- Adaptive polling adds state and complexity to the one tier that would otherwise be simple
Choose differently when. Per-tenant allocation is fine when tenants number in the thousands and each is commercially significant. The constraint comes from the tail's length, not from any dislike of reserved capacity.
Why it holds up over time. Making cost proportional to work rather than to configured entities is the economic principle under every generation of elastic infrastructure, and it will outlast the particular runtime that delivers it.
Lesson. When the business model is a long tail, the cost of an idle entity is an architectural constraint. Decide it before the design, because it eliminates options rather than tuning them.
Legibility
The author is not an engineer, and that is an architectural constraint rather than a UX preference.
ADR-16 · Every failure carries a classified, author-facing cause and a repair path
Status: Accepted · Shown on views: 21, 18, 05
A non-engineer's automation failed. What do they see, and what can they do about it?
Context. The product's premise is that someone who cannot read a stack trace can build and own this. That premise turns failure presentation into an architectural concern rather than a UI one: if the platform's internal states cannot be reduced to a sentence an author can act on, then either the states are wrong or the product is undeliverable. The default alternative — forwarding the provider's error — is worse than useless, because provider errors are written for the developer who called the API, and often describe conditions the author has no way to influence.
Decision. Every failure is classified into a small, stable, author-facing taxonomy, each class carrying a named next action. An unclassified error is never presented as the primary message, and an unclassified share above 2% of author-facing failures is treated as a defect in the taxonomy. A repair inbox groups parked and failed runs by cause with bulk retry, replay, discard and reconnect actions. The classification exists in the execution path, not in the presentation layer.
How it is realised on AWS. The runner writes a typed cause onto the ledger entry at the moment of failure, where the provider response and the step context are both still available; the run index projects it, and the repair inbox groups on it. Provider-specific error codes are mapped to platform classes in the connector manifest, so the mapping is versioned with the connector that knows about it. Redaction of credential material and author-declared sensitive fields happens before anything reaches run history.
| Option | Verdict | Reasoning |
|---|---|---|
| A typed taxonomy written at failure time, with a named action per class | Chosen | The only approach where the classification has access to the context needed to be accurate |
| Forward the provider's error message | Rejected | Cheapest and describes conditions in a vocabulary the author cannot act on, often about systems they do not administer |
| Classify in the presentation layer from stored error text | Rejected | Brittle string matching, far from the context, and silently wrong whenever a provider rewords a message |
| Route every failure to human support | Rejected | Does not scale past a tiny fraction of 1.8 billion monthly step attempts, and is slower than the author fixing it themselves |
What it buys
- An author can usually repair their own automation without support, which is the only workable path at this volume
- The unclassified rate becomes a measurable quality signal for the taxonomy itself, so gaps surface rather than accumulate
- Grouping by cause makes bulk repair possible: one reconnection fixes hundreds of parked runs at once
What it costs
- A mapping from provider errors to platform classes, per connector, which is ongoing work and another field that can be wrong
- The taxonomy must stay small to be legible and complete enough to cover reality, and those pull against each other continuously
- Redaction before storage means the raw provider response is not always available for operator debugging, which has its own cost
Choose differently when. Forwarding raw errors is correct when the users are developers who called the API themselves and have the context to interpret them. The taxonomy's cost is justified only by the non-engineer premise.
Why it holds up over time. The requirement that a system explain its failures in terms of the actions available to the person reading will outlast every interface technology, because it follows from who the user is rather than from how the message is rendered.
Lesson. If your users cannot act on your errors, your error handling is unfinished — and that is a property of where classification happens in the architecture, not of how the message is styled.
Every package used, in one table
Terms used in a specific sense in this package, where ordinary usage would be ambiguous.
| Package | What it is | What it does here | Considered instead |
|---|---|---|---|
| Effect key | A deterministic identifier derived from (run id, step id, logical attempt), passed to a provider as an idempotency key where the action's class says it will be honoured. | The mechanism that makes a retry harmless: the same key means the same operation, so a redelivery is recognised rather than repeated. | "Idempotency key", which in this package means the provider-side field the effect key is carried in, not the key itself. |
| Replay-safety class | A declared, versioned property of a connector action: idempotent (the provider honours a caller key), checkable (the effect can be read back by a natural key), or unsafe (neither). | The input from which every execution guarantee is derived, and a hard release gate on publishing an action. | "Retry policy", which in this package means the backoff schedule and attempt ceiling, a separate and much weaker idea. |
| Park | A normal run state with a release time, entered on a rate limit or an unresolvable ambiguity. It consumes no worker and does not spend the retry budget. | The state that lets the platform wait on somebody else's capacity without treating waiting as failure. | "Failed", which is terminal and consumes the budget; a parked run has neither property. |
| Retry | A further physical attempt at the same logical attempt, reusing the effect key. | Recovery from a fault, with no new effect intended. | "Replay", which is the opposite intention. |
| Replay | An author-initiated new logical attempt, with a new effect key, re-performing the effect deliberately from the retained trigger event. | Recovery of missed work, with a new effect explicitly intended and confirmed. | "Retry", which must never produce a second effect. |
| Step ledger | The append-only store of step attempts — intent before each outbound call, outcome after — keyed by run, step and attempt. | The platform's system of record for run state and the point a lost worker resumes from. Run status is a projection of it. | "Run history", which is the rebuildable author-facing projection, not the authority. |
| Trigger event log | The durable partitioned record of every accepted trigger event, retained for the replay window. | The platform's system of record for what happened, and the only thing a replay reads. | The providers, which remain the system of record for what is true. |
| Connection | One workspace's authorised grant into one provider account, carrying scopes, state and a reference to credential material held in custody. | The unit of credential custody, of provider rate limiting, and of cursor and subscription state. | "Workspace", which is the unit of tenancy, quota, billing and audit. |
| Quiet period | The interval after which a connection that has emitted nothing is considered unexpectedly silent, derived from its own observed history. | Turns an absence of data into a signal, which is the only way a dead subscription or a stuck cursor becomes visible. | A health check on the poller, which reports success while delivering nothing. |
| Unsafe action | An action for which the provider offers neither a caller-supplied idempotency key nor a reliable read-back. | The class for which the platform cannot deliver at-most-once visible effect, and says so rather than guessing. | "Dangerous" or "unsupported" — an unsafe action is fully supported; what is unavailable is the automatic resolution of an ambiguous outcome. |