An interviewer says — you run architecture for a consumer lending platform. The regulator expects evidence that your model is not producing disparate outcomes across ethnic groups. Legal tells you that you may not collect applicants' ethnicity. Where do you take this?
Show the full answer Hide the answer
What the interviewer is testing
Whether you treat "we are not allowed to collect it" as the end of the conversation. It is not. The obligation is to know the outcome distribution; the constraint is on using the attribute in the decision. Those are different data paths and a good architect separates them.
The clarifying questions that change the answer
- Which product and which jurisdiction? In US mortgage lending, demographic data is collected under HMDA and the question does not arise. For non-mortgage credit it generally is not collected, which is precisely why the CFPB published its 2014 proxy methodology white paper.
- Is the prohibition legal or a policy the firm chose? Many firms can lawfully collect self-reported data for monitoring only, with consent and purpose limitation, and have simply never asked.
- What decision would a finding change? If nobody will retrain or adjust a threshold, you are building a dashboard, not a control.
- What volume? This bounds what is detectable at all.
A strong answer's arc
Infer a probabilistic proxy in a separate environment. The documented approach is Bayesian Improved Surname Geocoding: combine the racial composition associated with a surname with that of the applicant's census geography into a probability across groups. The CFPB's white paper describes exactly this, and the code is public.
Then make the architecture carry the separation:
- A measurement enclave. Proxy probabilities are computed in a separate store, keyed to the decision record by a one-way identifier, with no path back into the feature store or the serving path. Access is a small named group, purpose-limited and logged.
- Aggregate only. Report group approval rates and score distributions, never per-applicant labels. A probabilistic proxy attached to an individual is both wrong often and a new sensitive attribute you now hold.
- Weight by probability rather than assigning a class. Thresholding at "most likely group" throws away the uncertainty and biases the disparity estimate, usually towards zero, because the misclassified applicants are concentrated in exactly the groups you are trying to compare.
- Choose self-reported data over a proxy whenever it is lawfully available, and reverse the choice only when collection is barred. The proxy costs accuracy and adds a sensitive dataset you did not previously hold; that trade-off is worth it only when the alternative is no measurement at all.
- State the power of the test. Detecting a five percentage point gap in approval rates around 50% needs roughly 1,600 decisions per group for a conventional 80% power, and the requirement scales with the inverse square of the gap: a two-point gap needs about six times as many. A monthly fairness dashboard on 300 applications is noise presented as assurance, and saying so is the difference between a senior answer and a compliant-sounding one.
Common weak answers
- "Remove the correlated features." Postcode, device, employer and income are correlated with everything; removing them degrades the model and leaves the gap. You cannot subtract your way to fairness.
- "Use a fairness library and pick demographic parity." Which definition to satisfy is a legal and commercial determination. Parity and calibration cannot both hold when base rates differ, so an engineer choosing silently has made a policy decision for the firm.
- "We cannot measure it, so the risk is accepted." Nobody senior can accept a risk whose size is unknown, and a regulator will treat the absence of measurement as the finding.
What a strong answer adds
The failure mode nobody mentions: the enclave is built, and eighteen months later a helpful engineer joins the proxy table to the feature store to "improve the model". Prevent it structurally — separate account, no network path, contract tests that fail if the proxy column appears in a training set, and the proxy pipeline owned by risk rather than by the modelling team. Add that the honest deliverable is a measured gap with a confidence interval and a decision about it, not a green tile.