practice

Negative Evidence Search

also called Failure-Side Evidence Review, Trouble-First Evaluation

Evaluating a technology from the record its troubled users left - issue labels, breaking-change notes, documented limitations - rather than from benchmarks, because that record predicts operating experience and benchmarks do not.

evaluationresearchupgradesissue trackerdue diligence

A team evaluating a datastore produces a benchmark report: throughput at three concurrency levels, p99 latency and a cost comparison. Everything in it is true and none of it predicts what the next 3 years will feel like, because a benchmark measures the system working and the next 3 years are decided by how it behaves when it is not.

The evidence that does predict that is already public and is written by people who had trouble. Issue trackers carry labels for the failure classes frequent enough to deserve a label. Release notes carry the breaking changes that set the real cost of staying current. Documentation carries the limitations the maintainers have decided not to fix. Operator blog posts titled after an incident carry the rest.

Negative evidence search is reading that side of the record first, deliberately and in a fixed order, before anybody runs a load test. It costs about half a day against a commitment measured in years.

Why it matters

Published technical material is dominated by two genres: marketing, which describes the system working, and success stories, which describe it working in a context that is not yours. Both are selected for. The failure record is the only large body of evidence that is not selected for a positive outcome, which is exactly why it is informative.

It also answers a specific question a benchmark cannot: what will the upgrades cost. A project that ships a breaking change in 3 of its last 4 majors has told you that staying current is a standing annual cost, and that number belongs in the business case next to the licence fee.

Implementation patterns

  • Issue labels first. Filter the tracker to the labels that matter — data loss, corruption, upgrade, rebalance, memory, deadlock — and read titles, not threads. The existence and size of a label is the signal.
  • One countable metric: the share of the top 20 issues by reactions that have been open more than a year. Above about half, plan to maintain a fork or to live with the behaviour.
  • The last 4 release notes' breaking-change sections, read as a series. This gives the upgrade cadence and the per-major cost directly.
  • The project's own documented limitations page, which is usually honest and usually unread.
  • Search for the operational verbs: "we migrated off", "postmortem", "why we stopped using", "upgrade from". Then weight by whether the author's scale resembles yours.
  • For a managed service, where the tracker is private, substitute the published status-page history over 18 months and two reference calls with operators at your scale.
  • Write the findings as 5 named risks with a mitigation each, so the output is comparable between candidates rather than a reading list.

Industry example

Reddit's 314-minute outage on 14 March 2023 turned on evidence that was already published. A Kubernetes upgrade from 1.23 to 1.24 removed a node role label; the cluster's Calico route reflectors were selected by that label, found no nodes, and pod networking collapsed. The label change was documented in the upgrade notes for that release, with migration guidance for workloads that selected on it. What was missing was not information but a reader with a reason to look — nobody owned the question "which of our selectors depend on strings this upgrade changes". The same asymmetry applies to any dependency: the negative evidence is usually there, and it is only useful to someone who went looking before committing.

Failure scenarios

  • Reading threads instead of counting labels, which takes 2 days and produces an anecdote.
  • Weighting by loudness. A heavily reacted issue from a workload 100 times yours may be irrelevant; one quiet issue matching your access pattern may be decisive.
  • Treating an old issue as evidence of a current defect, when it was fixed two majors ago.
  • Finding nothing and concluding the project is sound, when it is simply young and has few users. An empty failure record is a sign of low adoption, not of quality.
  • A report that lists risks with no mitigations, which gets read as obstruction and ignored.

Trade-offs

The practice is biased by construction: it will make every candidate look worse than its benchmark does, and read without discipline it argues for never adopting anything. It also costs calendar time at the point where a team is impatient to start building.

The compensating rule is to run it on every candidate including the incumbent and the do-nothing option. A failure record read for one option and not the others produces a worse decision than reading none of them, because the familiar choice gets credit for evidence nobody gathered.

When not to use it

Skip it for a reversible, contained choice: a library behind an interface in one service, a tool one team uses and can drop in an afternoon. The half-day exceeds the cost of being wrong.

Skip it also where the decision is already forced — a platform mandated by a parent company, a database the vendor's product requires. There the useful work is not evaluation but preparing for the failure classes you found, so convert the search into a runbook and a monitoring plan instead of a recommendation.

Prefer it strongly for anything that will hold state you cannot regenerate, where the failure classes are data loss and corruption and where the exit cost after 2 years is a migration project.

Interview question

Q: You have one week to evaluate a distributed database whose benchmarks are excellent and whose only production references are the vendor's own case studies. You cannot run a realistic load test in the time. What do you do with the week, and what single finding would stop the purchase?

What a strong answer covers: spending the first day on the failure record — issue labels for data loss, corruption and upgrade, the last 4 majors' breaking changes, the documented limitations · the countable signal of long-open high-reaction issues · finding an operator the vendor did not choose, and the questions that matter to them, which are upgrade pain and behaviour under partial failure · the one finding that should stop it, namely a data-loss or corruption class that is open, reproducible and matches your access pattern · and delivering 5 named risks with mitigations rather than a verdict.

Quick check

Quiz: Before reading a datastore's benchmarks, which sources tell you most about operating it? Its failure record: issue labels for data loss and upgrade problems, the last four majors' breaking changes, the documented limitations, and operator posts named after incidents.

Flashcard: What does an empty public failure record tell you about a project? — Usually that adoption is low, not that quality is high; a project with no complaints has no operators.