pattern

Decision Reconstruction Record

also called Decision Snapshot, Explainability Record

The set of facts captured at the moment an automated decision is made — inputs as scored, artefact versions and the rule trace — without which the decision cannot be explained or contested later.

gdprarticle-22automated-decisionsmodel-versioningaudit-trail

A lending platform declines applications with a model plus rules, and logs the outcome and the final score. A customer contests a decline from three months ago. The outcome and the score are the result; nothing recorded explains why, and nothing allows anyone to reproduce the decision as it was made.

Article 22 of the GDPR gives a right to human intervention and to contest a solely automated decision with legal or similarly significant effect, and Articles 13 to 15 require meaningful information about the logic involved. Human review is only meaningful if the human can reconstruct the decision — on the inputs as they were, under the artefacts as they were.

Why it matters

Source data mutates. An address is corrected, an income field updated, a credit file refreshed. A record that points at "the applicant's data" resolves to something different a month later, so the pointer is worthless exactly when the decision is challenged.

Model artefacts move too. Recomputing an explanation under the current model answers a different question and will sometimes contradict the original decision, which is worse than having no explanation because it implies the decision was wrong when it may simply have been different.

This is the one obligation that cannot be met retrospectively. Release history can partially recover which model was live; nothing can recover the input state.

Implementation patterns

  • Store the input feature vector as scored, by value, not by reference. This is the field that cannot be added later.
  • Record the identity and version of every participating artefact: model, feature pipeline, rule set, and any third-party score with the date it was obtained.
  • Capture the rule trace: which rules evaluated, in order, and which was decisive. In model-plus-rules systems the decline is usually a rule, and the score is a distraction.
  • Compute contributing factors at decision time and store them, rather than recomputing on demand under whatever is current.
  • Link the human-intervention outcome back to the original record, which makes the override rate a measurable property of the system and a useful quality signal.
  • Compute the explanation on demand from the record, rather than generating a narrative for every decision. Contested decisions are a tiny fraction of decisions made; the storage is the part that must be universal.
  • Keep the record for the limitation period of the relevant claim, which is usually longer than the retention anybody chose for the application's own data.

Industry example

Regulated credit decisioning is where the practice is most mature, driven by adverse-action explanation requirements in several jurisdictions and by model-risk governance expectations on banks. The EU AI Act, in force since August 2024 with obligations phasing in, adds record-keeping and logging requirements for high-risk systems including creditworthiness assessment, which pushes the same artefact from good practice towards a documented obligation. The engineering content of all these regimes converges on the same record, which is a useful argument when the requirement arrives from three directions at once.

Failure scenarios

  • Inputs by reference, so the reconstruction uses corrected data and produces a different outcome than the one being contested.
  • A recomputed explanation that contradicts the decision, handed to a complainant or a regulator.
  • No rule trace, so a rule-driven decline is explained in terms of a model score that was not decisive.
  • Model version absent, leaving the team to infer it from deployment logs, which is both slow and arguable.
  • A third-party score with no retrieval date, where the bureau's own data has since changed and the vendor cannot reproduce what it returned.
  • Records retained for less time than the right to contest, which is the quiet failure: the design was correct and the retention policy deleted the evidence.

Trade-offs

Writing a record on every decision costs storage and a write on the critical path. The sizing usually ends the argument: a few kilobytes per decision means a million decisions a month is a few gigabytes, trivial against the cost of being unable to explain one decision. The genuine costs are elsewhere — a versioning discipline across models, pipelines and rule sets, and the privacy question of holding a snapshot of personal data for years, which must itself be justified, minimised and access-controlled.

When not to use it

Where a decision has no legal or similarly significant effect on a person — a content ranking, a layout variant, a search order — the regime does not engage and this machinery is disproportionate. The test is the effect on the individual, not the sophistication of the model: a simple rule with a significant effect needs the full record, and a complex model with a trivial effect does not. For a decision that is only ever advisory to a human who makes the real call, the record of the human's reasoning matters more than the model's.

Interview question

Q: A regulator asks you to explain a specific automated decline made four months ago. Your system logs the outcome, the score and a timestamp. Tell me what you can and cannot say, what you would do in the next two weeks, and what you would say to the regulator about the gap.

What a strong answer covers: stating plainly that the decision cannot be reconstructed, and why input mutation makes this irreparable for past decisions · refusing to recompute under the current model as a substitute, and explaining why that is worse than admitting the gap · the record design, with inputs by value as the critical field · rule trace versus score in a hybrid system · retention aligned to the contest period rather than to the application's data policy · and a remediation plan that starts recording immediately while being honest about the historical window.

Quick check

Quiz: Why is storing a reference to the applicant's data insufficient for a decision record? — Because the referenced rows mutate, so the reconstruction uses different inputs than the decision did, and the resulting explanation may contradict the outcome being contested.

Flashcard: In a model-plus-rules decisioning system, which part of the record is most often missing and most often decisive? — The rule trace: the decline is usually caused by a rule firing, while the logged model score is a distraction.