advanced 3 min answer

Review this. A lending platform declines applications using a model plus rules, and logs the decision outcome and the final score. Legal says the design satisfies the GDPR's Article 22 requirements on automated decision-making because a human can review any decline on request. What is wrong?

gdprarticle-22automated-decisionsexplainabilityaudit-trail
Show the full answer Hide the answer

What is actually required

Two separate things the design conflates. Article 22 gives a right to human intervention and to contest a decision, and Articles 13 to 15 require meaningful information about the logic involved. A human review is only meaningful if the human can reconstruct the decision that was made — on the inputs as they were, under the model and rules as they were, at the time it was made.

The current design cannot do that. An outcome and a score are the result. Nothing recorded explains why, and nothing allows anyone to determine whether the decision would be the same today for a reason other than drift.

What I would change

Record a decision reconstruction record at the moment of the decision:

  • The input feature vector as scored, not a reference to the source rows. Source data mutates: an address is corrected, an income field is updated, a credit file changes. A pointer to "the applicant's data" resolves to something different a month later, which makes the record worthless precisely when it is contested.
  • The identity and version of every artefact that participated: model version, feature-pipeline version, rule-set version, and the version of any third-party score consumed, with the date it was obtained.
  • The rule trace: which rules fired, in order, and which one was decisive. In a model-plus-rules system the decline is usually a rule, and the score is a distraction.
  • The contributing factors, computed and stored at decision time rather than recomputed later. Recomputation under a newer model answers a different question and will sometimes contradict the original decision.
  • The human-intervention outcome, when it happens, linked to the original record, so the override rate becomes a measurable property of the system.

The one change that matters most

Store the inputs as scored. Everything else can be partially recovered from release history; the input state cannot be recovered at all once the source rows have moved on. Teams discover this during their first contested decision, typically months after go-live, and the honest answer at that point is that the decision cannot be explained.

What I would leave alone

The model plus rules structure is fine and often preferable to a pure model, because the decisive rule is explainable in a sentence. Human review on request is also correct as far as it goes. And do not build a real-time explanation service for every decision: store the record and compute the explanation on demand, because the number of contested decisions is a tiny fraction of the number made, while the storage is the part that cannot be added retrospectively.

How I would argue this in review

With a dated example. Take a decline from three months ago and ask the team to demonstrate why it was declined. If the answer requires reasoning about what the data probably looked like, the record is insufficient, and the demonstration is more persuasive than any citation. Then size it: a few kilobytes per decision, so a million decisions a month is a few gigabytes — trivial storage against the cost of being unable to explain a single decision to a regulator or a court.

When this is more than is needed

Where a decision has no legal or similarly significant effect on a person — a content ranking, a layout choice — Article 22's regime does not engage and this machinery is disproportionate. The test is the effect on the individual, not the sophistication of the model. A simple rule with a significant effect needs the full record; a complex model with a trivial effect does not.