practice

Spike Budget

also called Price of Information, Information Purchase Test

The arithmetic that decides whether to buy evidence before committing - comparing the full cost of an investigation against the probability of choosing wrong multiplied by what correcting it later would cost.

spikeevaluationestimationreversibilitydecision-under-uncertainty

"Let us run a spike" sounds free and is not. Two engineers for two weeks is 4 engineer-weeks, roughly 160 hours, about $16k to $24k fully loaded in 2026 in a high-cost market. The number teams omit is the calendar: two weeks of delay on a feature worth roughly $200k a quarter is about $30k of deferred value. The spike costs about $50k, and the decision to run one should be made against that figure.

The test is one line:

price of investigation  <  P(wrong without it) × (cost of correcting later − cost of correcting now)

At P(wrong) near 0.4 — two plausible options, no first-hand evidence — against a $150k migration, the expected avoided cost is $60k and the spike is worth buying. If the team already runs one of the options in production, P(wrong) falls to roughly 0.1, the expected avoided cost is $15k, and the correct move is to decide today and spend the two weeks shipping.

Why it matters

Both defaults are expensive. "Always evaluate first" treats information as free and cannot survive 20 decisions a quarter. "Just pick one" is right most of the time and fails precisely on the one-way doors where correction costs a quarter. The arithmetic separates those cases and moves the argument from conviction to two numbers.

Implementation patterns

  • Price the delay as well as the hours. It is usually the larger term and it is the one that makes a six-week evaluation indefensible: at roughly $150k all-in it costs as much as being wrong.
  • Write the flip condition before starting. "If the managed service cannot sustain 40M updates a day inside the ingest window at list price, we self-host." An investigation with no observation that would change the choice is not an experiment.
  • Timebox it and treat an overrun as data. A two-week spike that runs five weeks has told you the operational cost of that option is higher than the proposal assumed.
  • Compare against buying reversibility. Putting the choice behind an interface can cut the correction cost from $150k to $40k, at which point no spike clears the bar.

Industry example

The cheapest information is bought in production rather than in a laboratory. Take the archetype of a travel-search platform aggregating supplier prices, the business shape Expedia operates in: the choice between a managed search service and running the engine is not settled by a benchmark, because the question is behaviour under 40 million price updates a day with a nightly reindex. The cheapest move routes one supplier feed - a few percent of volume - through the managed option for two weeks behind the same interface. That answers the ingest-window question with real data and stays a shipped slice rather than throwaway work. The transferable rule is the ordering: buy cheap information first, prefer information that doubles as delivery, and spend real money only where the correction cost is large and the prior is near a coin flip.

Failure scenarios

  • The spike that becomes a project, with no stop condition.
  • The spike that cannot change the answer, run to create the appearance of diligence. This is the $50k purchase of comfort.
  • P(wrong) assumed rather than argued, which is where the whole calculation lives: two people with the same costs and different priors are arguing about the prior.

Trade-offs

Buying information pays in calendar time, engineer attention and the morale cost of work that ships nothing. Buying reversibility pays in abstraction: an interface with two implementations is more code, leaks the weaker option's limitations, and tempts the team to never commit. They are substitutes, and reversibility is usually cheaper — so the first question is how expensive being wrong would be, not how uncertain you feel.

When not to use it

Skip the arithmetic where the correction cost is small; the calculation costs more than the mistake. Skip it where the uncertainty is organisational rather than technical, since whether a team still owns a system in two years is not resolvable by a spike. And where the information is nearly free - a half-day reading two postmortems from operators at your scale - skip the test and read them.

Interview question

Q: "You want two weeks to evaluate two datastores. Your director says pick one today. Make the case with numbers, and then tell me under what conditions you would agree with the director."

What a strong answer covers: the full price including delay; an explicit prior for choosing wrong and where it comes from; the correction cost in engineer-weeks; the flip condition the spike would test; and the agreement case — a reversible choice, a prior near 0.1, or a cheaper substitute such as one reference conversation.

Quick check

Quiz: A two-week spike costs $20k in engineer time and delays a feature worth $200k a quarter. What is the real price, and what else must you know? About $50k once the delay is counted; you also need the probability of choosing wrong and the cost of correcting later, and you buy the spike when its price is below their product.

Flashcard: What is the cheapest substitute for a spike? Buying reversibility — putting the choice behind an interface — which often cuts the correction cost enough that no investigation clears the bar.