Fairness & Bias intermediate 8 min read 6 flashcards

Discrimination in LLM-Mediated Decisions

The correspondence-study design adapted to prompted models, what Anthropic's discrim-eval and the realistic-context follow-up measured, and why an anti-bias prompt that passes a clean evaluation fails once the prompt looks like production.

The useful measurement design here comes from labour economics, not from machine learning. Send identical résumés to real job ads, vary only the name, count the callbacks. White-sounding names received about 50 percent more callbacks than African American-sounding ones (Bertrand and Mullainathan, 2004, Are Emily and Greg More Employable Than Lakisha and Jamal?, American Economic Review). A language model reading a prompt is a far easier subject than a hiring manager: you can run the paired comparison millions of times, hold everything else byte-identical, and get the model's own probability for each outcome.

This matters because an LLM in a triage, screening or eligibility path has no confusion matrix to slice. There are no scores, often no labels, and the input is a document rather than a feature vector. Perturbation is what replaces the sliced metric.

Perturbation audits

The design: construct decision scenarios, vary the demographic signal while holding the substance fixed, and measure the change in the model's favourable-decision probability. Anthropic's discrim-eval generated 70 decision scenarios spanning finance, employment, housing and public services, systematically varied age, race and gender in each, and reported a discrimination score capturing how much more likely a favourable decision becomes for one demographic than another. Claude 2.0 showed both positive and negative discrimination in some settings with no intervention, and prompt-level interventions, stating that demographics must not affect the decision and asking the model to reason about avoiding it, reduced both substantially (Tamkin et al., 2023, Evaluating and Mitigating Discrimination in Language Model Decisions, arXiv:2312.03689; dataset).

Two properties of the design are doing the work. It is paired, so between-case variance cancels and the comparison is far tighter than a between-subject slice of production traffic. And it is generative, so coverage is a choice rather than an accident of what traffic arrived.

The clean evaluation is the optimistic one

Then the context gets realistic, and the result moves. Adding a company name, culture text lifted from a public careers page, and a selective hiring constraint to otherwise identical prompts induced racial and gender differences in interview rates of up to 12 percent, in a setting where the anti-bias prompting that worked on the stripped-down version no longer did. An intervention at the representation level, identifying directions in the model's activations that encode the sensitive attribute and neutralising them, held up across every scenario tested, typically under 1 percent residual bias and never above 2.5 percent, generalising from directions found on a simple synthetic dataset (Karvonen and Marks, 2025, Robustly Improving LLM Fairness in Realistic Settings via Interpretability, arXiv:2506.10922). See activation steering and representation engineering for the mechanism.

The transferable lesson is about evaluation, not about steering vectors. A mitigation validated on a minimal prompt has been validated on a distribution you do not serve.

Designing one that means something

The unit of analysis is the (case, demographic) pair, and the decisions that determine whether the result is interpretable are made before any model runs.

Explicit or implicit signal. Stating "the applicant is a 45-year-old Black woman" measures the model's response to a stated attribute. Swapping Lakisha for Emily measures its response to a correlated surface cue, which is both more realistic and confounded, because names carry regional, class and generational information too. Report which you did; they are different experiments with different relationships to deployment.

Repetition, because the system is stochastic. At non-zero temperature a single pair tells you almost nothing. Score the token probabilities where you can, and where the decision only exists as sampled text, repeat each pair enough times that the per-pair estimate has a usable interval before you aggregate.

Order and format effects. Comparative prompts ("which of these two candidates") inherit position bias from the model, so counterbalance the order or you will measure it instead of discrimination (see prompt format sensitivity).

When it breaks

Generated scenarios are not your traffic. A public eval tells you about a model. It does not tell you about your prompt, your retrieval context, your rubric or your document distribution. The audit that supports a deployment claim has to run on your own prompts.

Favourable-direction bias is still a legal problem. A model that tilts toward a protected group is disparate treatment, not a win, and an audit reporting only the worst-off group will not see it (see disparate impact and the legal frame).

Prompt-level mitigations are brittle in a specific direction. They degrade as the surrounding context grows richer, which is the direction production moves in. Re-measure after any prompt or context change, and do not treat a one-off result as a property of the system.

The model is one stage. If a human reviews the model's shortlist, the quantity that matters is the disparity in the final decision. A debiased ranker feeding a reviewer who trusts it selectively can produce a worse outcome than no model, and only end-to-end measurement will show it (see appropriate reliance and its two failures for why that reliance pattern is the norm rather than the exception).

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Bertrand and Mullainathan, 2004, Are Emily and Greg More Employable Than Lakisha and Jamal?, American Economic Review scholar.harvard.edu
  2. Tamkin et al., 2023, Evaluating and Mitigating Discrimination in Language Model Decisions, arXiv:2312.03689 arxiv.org
  3. dataset huggingface.co
  4. Karvonen and Marks, 2025, Robustly Improving LLM Fairness in Realistic Settings via Interpretability, arXiv:2506.10922 arxiv.org
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track