Developer Environments
also called Developer Environment Tiers, Pre-Production Tiering
The set of places an engineer can run and exercise the system before production, and the deliberate decision about which classes of defect each tier is allowed to be unable to find.
A team has three places to run its service: a laptop, an environment created per pull request, and a shared pre-production estate. Each costs differently and each is blind to a different class of defect. The failure is rarely that a tier is missing; it is that nobody wrote down which defects each tier can catch, so a change that passed all three is believed safe when the classes that cause incidents were never testable at all.
Treat the tiers as a portfolio: assign every defect class to the cheapest tier that can detect it, then say out loud which classes no tier catches. Those are not gaps to close with another environment - they are the reason progressive delivery exists.
Why it matters
Feedback time is the variable. A defect caught by a unit test costs seconds; caught in a shared environment it costs the queue plus a deploy plus someone else's debugging; caught in production it costs an incident. Each tier upward is roughly an order of magnitude more expensive per defect found.
The second reason is honest coverage. A per-change environment with one user and 200 rows of seed data cannot exhibit lock contention, pool exhaustion, a plan flip at real cardinality or a concurrent-write race.
Implementation patterns
- Tier 1, the laptop stack: unit and contract tests, the service plus its own datastore. Target under 60 seconds for a loop run dozens of times an hour.
- Tier 2, an environment per change, destroyed on merge: wiring, migrations, configuration, contracts, deploy mechanics. Provision in 5 to 15 minutes, with an expiry so cost cannot accumulate.
- Tier 3, one shared estate with production-shaped data volume: plan changes, index behaviour, pool sizing, cache behaviour under churn. Expensive, so there is one and it is not per team.
- Write the matrix down: defect class against tier, with explicit "not detectable here" cells.
- Give tier 2 a second concurrent writer and a crude load generator. Two writers at 200 requests per second finds much of the contention class for almost nothing, and seed production-shaped cardinality rather than production data, because plan flips follow statistics rather than content.
Industry example
Cloud-hosted developer environments are a product category now, and vendor pricing states the economics: the machine is billed hourly, so the stop policy dominates. A 4 vCPU and 16 GiB workspace at roughly $0.15 to $0.20 an hour at 2025 list prices costs about $25 to $30 a month running only in working hours and $110 to $150 if left running, since a month holds about 160 working hours against 730. Across 400 engineers that is roughly $120k against $600k a year. The same logic applies inside tier 1: IDE vendors now publish shared index artefacts so the project index is built once centrally rather than by every engineer.
Failure scenarios
- The contention blind spot shipped as confidence. Tier 2 and tier 3 pass; production takes lock waits, because neither tier ran two writers at volume.
- Tier 3 becomes a queue, so changes are batched into it and the batch destroys the attribution the tier existed to give.
- Tier 2 without expiry, which is how a successful pilot becomes a tripled cloud bill with no owner.
- Mock-shaped success: tier 1 passes because the mock answers the way its author expected.
Trade-offs
| Choose | Gains | Pays |
|---|---|---|
| Heavy tier 1 and light tier 3 | Fastest feedback and lowest spend | Nowhere to find concurrency and scale defects |
| Per-change tier 2 for everything | Clean attribution and no shared queue | Spend scales with pull-request volume |
| One production-shaped tier 3 | Catches the expensive classes before users do | A queue, drift and a large standing bill |
When not to use it
One service, one datastore and five engineers needs one tier and a staging deploy, since a per-change environment pipeline costs weeks against a problem that was not hurting.
Skip tier 3 where production can be exercised safely instead: flags, shadow traffic and canaries catch the scale-and-concurrency class with real data. Prefer production-based detection over a fourth environment whenever a change can be exposed to a small share of traffic and reversed in seconds.
Interview question
Q: Local setup, a per-pull-request environment and a shared staging all passed a payments change, and it still caused a 9-second checkout latency incident from lock waits on one table. The team wants a fourth environment. What do you tell them?
What a strong answer covers: that the defect class needs concurrency and production-shaped cardinality, so no tier of that shape could catch it · the two cheap additions that would have, a concurrent writer and real row counts in tier 2 plus a migration lock check in the pipeline · that the residue belongs to progressive delivery, so the money goes on a canary with a latency-based abort · and writing the defect-class matrix.
Quick check
Quiz: Which defect classes can a single-user per-pull-request environment never exhibit? Lock contention, pool exhaustion, plan flips at real cardinality and concurrent-write races - all need concurrency and production-shaped data.
Flashcard: What decides how many developer environment tiers you need? — Not team size but the defect classes you must find before production; a tier earns its cost only by catching a class the cheaper one cannot.