The AI Learning Hub Journal

Fair Lending and Proxy Discrimination

The model does not see it — and reconstructs it anywaya model with enough correlated features rebuilds a characteristic nobody supplied to it, and nobody chose thatEXCLUDED FROM THE INPUTS race · ethnicity · gender postcode and geography employer education institution occupation category channel and device language of application merchant and transaction patterns alternative data for thin-file applicants the model facially neutral materially worse outcomes for a protected groupreconstructed without anyone choosing it — richer feature sets increase proxy capacity THE SUBTLER PROBLEM IS THE LABELa credit model learns from outcomes observed only onapplicants approved under previous policy — so historicdecisions are baked into the target the new model istrained to reproduce; reject inference estimates theunobserved outcome, which is an estimate, not a fix TESTING FOR DISPARATE OUTCOMESyou cannot test what you do not measure — some US mortgagecontexts require collecting applicant demographic data whileother regimes restrict collecting or inferring it at allestimated group membership carries error into the resulttest pricing, limits, terms, overrides and error rates too WHERE A DISPARITY APPEARS IN US LENDING — WHAT THE ECOA AND FAIR HOUSING ACT STRUCTURE ASKS NEXT a substantial legitimate business need? a less discriminatory alternative that meets it? the record of the search and the options set aside fairness metrics conflict mathematically, so choosing between them is a senior policy decision · testing standards for LLM components in a decision chain are unsettled “THE MODEL DOES NOT SEE RACE” IS A STATEMENT ABOUT INPUTS — THE LAW ASKS ABOUT OUTCOMES in US lending, discriminatory effect can support a challenge without any showing of intent
Fair lending exposure runs on outcomes, so excluding the variable is not by itself a defence.

Removing the Variable Does Not Remove the Effect

"The model does not see race" is a statement about inputs, and fair lending law is largely concerned with outcomes. In US lending, the Equal Credit Opportunity Act and the Fair Housing Act support challenges based on discriminatory effect and not only on intent, so a facially neutral model that produces materially worse outcomes for a protected group is exposed regardless of what it was fed. Other jurisdictions reach similar territory through equality and consumer protection law. The mechanism is proxying: a model with enough correlated features can reconstruct a characteristic it was never given, and it does so without anyone choosing that. More features and richer data increase proxy capacity rather than reducing it.

  • Fair lending exposure runs on outcomes, so excluding the variable is not by itself a defence
  • In US lending, discriminatory effect can support a challenge without any showing of intent
  • Correlated features let a model reconstruct a characteristic nobody supplied to it
  • Richer feature sets increase proxying capacity — more data is not automatically fairer

Where Proxies Actually Come From

The usual suspects are geography and postcode, which in many countries encode residential segregation directly. Then education institution, employer, occupation category, channel and device, language of application, and transaction or merchant patterns that differ systematically by community. Alternative data brought in to serve thin-file applicants can help genuine inclusion and can also import new proxies at the same time. The subtler problem is the label. A credit model learns from outcomes observed only on applicants who were approved under previous policy, so historic decisions are baked into the target the new model is trained to reproduce. Reject inference attempts to correct for that, but what it produces is an estimate of an unobserved outcome rather than a fix.

  • Postcode, employer, education, device, channel and merchant patterns are the common proxy carriers
  • Alternative data can widen access and import new proxies in the same step
  • The training label reflects who was approved before, so old policy is inherited by the new model
  • Reject inference partially addresses the unobserved-outcome problem; it does not eliminate it

Testing for Disparate Outcomes

You cannot test what you do not measure, and here the law pulls in two directions. Some US mortgage contexts require collection of applicant demographic information; in several other jurisdictions data protection rules restrict collecting or inferring the same characteristics, which leaves firms trying to test fairness without the data that makes testing possible. Statistical proxy methods used to estimate group membership are themselves estimates and carry their own error, which propagates into the test result. Whatever the approach, test more than approval rates: examine pricing, assigned limits, terms offered, override patterns and error rates by group, and record the methodology so a supervisor can see what you did and what you could not do.

  • Testing requires demographic data that some regimes mandate and others restrict — say which applies to you
  • Estimated group membership carries error that flows straight into the fairness result
  • Test approvals, pricing, limits, terms, overrides and error rates — not approval rates alone
  • Document the method and its limits; an undocumented test is not evidence to a supervisor

The Less-Discriminatory-Alternative Question, and Honest Limits

Where a disparity appears in US lending, the ECOA and Fair Housing Act framework runs a burden-shifting structure: whether the practice serves a substantial legitimate business need, and whether a less discriminatory alternative would meet that need comparably. What the framework asks for, then, is a search for such an alternative and a record of it, including the options considered and why each was set aside. Other jurisdictions reach related questions through equality and consumer law. Two caveats belong alongside all of this. Fairness metrics conflict mathematically, so no model satisfies every reasonable definition at once and the choice of metric is a policy decision recorded at a senior level. And practice for language-model components inside a decision chain is genuinely unsettled.

  • In US lending, the ECOA and Fair Housing Act structure asks about business necessity and less discriminatory alternatives
  • What that framework asks for is the search and the record of it, including the options set aside and why
  • Fairness definitions conflict mathematically; choosing between them is a senior policy decision, not a technical one
  • Testing standards for LLM components inside a decision chain are not settled — say so rather than imply rigour

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.