Fair Lending and Proxy Discrimination
Removing the Variable Does Not Remove the Effect
"The model does not see race" is a statement about inputs, and fair lending law is largely concerned with outcomes. In US lending, the Equal Credit Opportunity Act and the Fair Housing Act support challenges based on discriminatory effect and not only on intent, so a facially neutral model that produces materially worse outcomes for a protected group is exposed regardless of what it was fed. Other jurisdictions reach similar territory through equality and consumer protection law. The mechanism is proxying: a model with enough correlated features can reconstruct a characteristic it was never given, and it does so without anyone choosing that. More features and richer data increase proxy capacity rather than reducing it.
- Fair lending exposure runs on outcomes, so excluding the variable is not by itself a defence
- In US lending, discriminatory effect can support a challenge without any showing of intent
- Correlated features let a model reconstruct a characteristic nobody supplied to it
- Richer feature sets increase proxying capacity — more data is not automatically fairer
Where Proxies Actually Come From
The usual suspects are geography and postcode, which in many countries encode residential segregation directly. Then education institution, employer, occupation category, channel and device, language of application, and transaction or merchant patterns that differ systematically by community. Alternative data brought in to serve thin-file applicants can help genuine inclusion and can also import new proxies at the same time. The subtler problem is the label. A credit model learns from outcomes observed only on applicants who were approved under previous policy, so historic decisions are baked into the target the new model is trained to reproduce. Reject inference attempts to correct for that, but what it produces is an estimate of an unobserved outcome rather than a fix.
- Postcode, employer, education, device, channel and merchant patterns are the common proxy carriers
- Alternative data can widen access and import new proxies in the same step
- The training label reflects who was approved before, so old policy is inherited by the new model
- Reject inference partially addresses the unobserved-outcome problem; it does not eliminate it
Testing for Disparate Outcomes
You cannot test what you do not measure, and here the law pulls in two directions. Some US mortgage contexts require collection of applicant demographic information; in several other jurisdictions data protection rules restrict collecting or inferring the same characteristics, which leaves firms trying to test fairness without the data that makes testing possible. Statistical proxy methods used to estimate group membership are themselves estimates and carry their own error, which propagates into the test result. Whatever the approach, test more than approval rates: examine pricing, assigned limits, terms offered, override patterns and error rates by group, and record the methodology so a supervisor can see what you did and what you could not do.
- Testing requires demographic data that some regimes mandate and others restrict — say which applies to you
- Estimated group membership carries error that flows straight into the fairness result
- Test approvals, pricing, limits, terms, overrides and error rates — not approval rates alone
- Document the method and its limits; an undocumented test is not evidence to a supervisor
The Less-Discriminatory-Alternative Question, and Honest Limits
Where a disparity appears in US lending, the ECOA and Fair Housing Act framework runs a burden-shifting structure: whether the practice serves a substantial legitimate business need, and whether a less discriminatory alternative would meet that need comparably. What the framework asks for, then, is a search for such an alternative and a record of it, including the options considered and why each was set aside. Other jurisdictions reach related questions through equality and consumer law. Two caveats belong alongside all of this. Fairness metrics conflict mathematically, so no model satisfies every reasonable definition at once and the choice of metric is a policy decision recorded at a senior level. And practice for language-model components inside a decision chain is genuinely unsettled.
- In US lending, the ECOA and Fair Housing Act structure asks about business necessity and less discriminatory alternatives
- What that framework asks for is the search and the record of it, including the options set aside and why
- Fairness definitions conflict mathematically; choosing between them is a senior policy decision, not a technical one
- Testing standards for LLM components inside a decision chain are not settled — say so rather than imply rigour
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.