Product teams need a Kafka topic and an object-storage bucket per service with quotas and retention that change over time. Should the platform expose a synchronous provisioning call that returns when the resource exists or a declarative resource reconciled by a controller or a pull request into a shared Terraform repository?
Show the full answer Hide the answer
The deciding property
Not how developers feel about YAML. The settling fact is that these resources have a lifetime: quotas, retention and access policies change after creation, and they drift — someone widens a policy by hand, a quota is raised during an incident, a cloud default changes under you. A provisioning interface that only creates has to grow a second interface for every subsequent change, and a third to detect what no longer matches.
The second fact is that creation is slow and partially fails. A topic plus a bucket plus a role plus a policy is four external calls over minutes, and failure halfway leaves a real resource that nobody is tracking.
Why a reconciled resource
A declarative record of intent plus a controller that repeatedly compares intent with reality gives three things for one mechanism:
- Retries are free and idempotent. The controller's job is convergence, so a timeout is not a special case and a half-built resource is just a state to continue from. This is the difference between a 40-minute workflow that fails 15% of the time and one that finishes eventually without a human.
- Drift is corrected rather than discovered. The same loop that creates also re-asserts. A hand-widened policy is reverted within one reconcile interval, and the platform can report how often that happens, which is a governance answer nobody had before.
- The consumer contract stays small. Teams declare the end state; the platform keeps freedom to change how it gets there, including changing cloud provider, because the intent record is the interface and the implementation is behind it.
What it costs: there is no synchronous success and no single place where failure surfaces, so status must be first class — conditions on the resource, a clear "not ready and why" state, and events the owning team can read without platform help. Teams used to an error code find "eventually" harder to reason about, and that is real.
Why the other options fail
- A synchronous provisioning call is right when the caller genuinely cannot continue without the resource and the work fits in one request — an ephemeral namespace inside a 90-second CI job, for example. For a four-call multi-minute workflow it converts every partial failure into an orphaned resource and a caller-side retry that duplicates work.
- The Terraform pull request is right more often than platform teams admit: for perhaps ten teams and twenty resources a month, a reviewed module in one repository plus a README beats a controller you must now operate. It fails here because human-paced review does not correct drift, and because the repository becomes the queue the platform was built to remove.
- A request queue with approval is the model being escaped. Automation behind a human gate keeps the queue's latency and adds the automation's maintenance. The exception is a genuinely irreversible or expensive action, where the approval is the product.
What would flip the decision
| If this changes | Choose | Because |
|---|---|---|
| The caller must know synchronously and the work is one short call | Synchronous API | reconciliation gives no answer in time |
| The estate is small and stable and nothing drifts | Terraform module and review | a controller is an operational dependency you do not need |
| Resources are created once per year and cost thousands | Approval with automation | the review is the value |
Common weak answers
"Expose both and let teams pick" doubles the surface and the drift. "Declarative is best practice" without naming drift and partial failure is the answer that gets built and then abandoned when nobody can explain why a bucket is missing. The rule to carry: reconcile when the resource has a life after creation; call directly when it does not.