What Regression Analysis Can and Cannot Prove in Fair Lending
Published 18 August 2026 · Quoted from the OCC Comptroller’s Handbook, Fair Lending, Version 1.0 (January 2023, as amended 14 July 2025) and the FFIEC Interagency Fair Lending Examination Procedures (August 2009). Both linked at the foot of the page.
The short answer. A statistical model can establish that a difference in outcome survives the controls you gave it, and whether that difference is unlikely to have arisen by chance. That is real and useful. What it cannot establish is that no legitimate factor you left out explains the result — because a model measures what is unexplained by the variables it was given, which is not the same thing as what is unexplainable.
The regulators are less credulous about this than most vendor marketing. The OCC Comptroller’s Handbook publishes a specific three-part test for what to do when a factor is missing from a model, and it is worth knowing before you rely on your own results.
First, the word itself
The current OCC Comptroller’s Handbook on fair lending does not use the word “regression” anywhere. It says statistical modeling. The FFIEC Interagency Fair Lending Examination Procedures say statistical analysis, naming it alongside comparative file review as the approaches available.
Regression is the usual implementation of statistical modeling in this setting, so these are not different concepts — but it is worth knowing which vocabulary belongs to whom. A vendor feature list that advertises “regression analysis” and an examination procedure that describes “statistical modeling” are talking about the same territory. We say regression on our own product pages, and it is the industry's word rather than the examiners'.
What a model can establish
Given a set of legitimate underwriting or pricing factors, a model can tell you that a difference in outcome between a prohibited basis group and a control group persists after those factors are accounted for, and can attach a significance level to that difference. That is substantially more than a raw approval-rate gap tells you, because it separates “these groups had different outcomes” from “these groups had different outcomes even among applicants who looked similar on the things we measured.”
It can also, usefully, produce application-level residuals — a per-application measure of how far each decision sits from what the model predicted. Those are the files worth reading, which is the practical bridge from a model to a review.
What a model cannot establish, and the regulator's own test
The limitation is the omitted variable. If a factor genuinely drove the decisions, is correlated with the prohibited basis characteristic, and is not in your model, the model will attribute its effect to whatever is in the model — potentially to the prohibited basis variable. This is not a defect of any particular software; it is a property of the method.
The OCC handbook addresses it directly, and this is the passage worth having in front of you when you read your own output:
- are highly correlated with the PB characteristic.
- themselves have a strong effect on the outcome (e.g., underwriting and pricing).
- have a systematic effect.”
That is a usable framework in both directions. If a factor you omitted fails all three tests — weakly correlated with the protected characteristic, weak effect on outcome, idiosyncratic rather than systematic — its absence probably does not undermine the result. If it passes all three, your model is telling you less than it appears to.
Three further things a model does not do, worth being explicit about because they are routinely assumed:
- It does not establish causation. An unexplained disparity is an unexplained disparity.
- It does not distinguish policy from discretion. A model cannot tell you whether a pattern comes from a written rule applied consistently or from individual judgement exercised inconsistently, and the remediation for those two is different.
- It does not tell you which files to fix. Residuals point at candidates; reading the files is what identifies whether anything actually went wrong.
The file review is not a rebuttal, and not a replication
This is where institutions most often get the relationship wrong in both directions, and the handbook is precise about it:
- determine whether there are potentially significant data errors or any bank policies that are incorrectly accounted for in the model.
- determine the robustness of the model by identifying factors that were not, but could be, incorporated into the statistical model.”
So a file review is a test of the model, not a second opinion competing with it. And the handbook constrains what a file review is allowed to change:
“It is important not to dismiss statistical results based simply on findings from a manual file review.” OCC Comptroller’s Handbook, Fair Lending, Version 1.0, p.19
Read those together and the asymmetry is deliberate. You may not wave away a statistical disparity by explaining a handful of files individually. If you want a factor added to the model, it has to come from your written policy and it has to operate systematically — not be a post-hoc rationalisation assembled from the exceptions. That is a governance requirement as much as a statistical one: the factors that can defend you are the ones you wrote down before you needed them.
Statistical significance is not a finding
Significance is a statement about how unlikely a result is under the model, not a legal conclusion and not a statement of cause. A significant result is a reason to investigate: check the data, test for omitted factors, read the files. Treating a threshold as the answer is how a self-assessment generates either a false alarm or false comfort, and a compliance officer who has been on the receiving end of an overconfident model already knows this.
The corollary matters too: a non-significant result is not a clean bill of health. Low volume in a product produces wide confidence intervals, and “we could not detect a disparity” is a different claim from “there is no disparity.” Where volume is thin, the file-based analysis is not a lesser substitute — it is the appropriate method, which is the point covered in our piece on comparative file review versus matched-pair testing.
If you are citing this booklet, look at the page. And note that the removal is an edit to the OCC’s examination booklet: ECOA, Regulation B and the Fair Housing Act are the underlying law, the FFIEC Interagency Procedures are a separate document, and institutions supervised by other agencies are examined under those agencies’ materials. Anything turning on this is a question for your counsel, not for a software vendor.
If you are running this yourself
- Write down your factor list and where each factor comes from in policy, before you run anything. That document is what lets you add a factor later without it looking invented.
- Test your own omitted variables against the three-part test above. If a factor you excluded is correlated with the protected characteristic, strong on outcome, and systematic, you have a model problem rather than a finding.
- Use residuals to select files, then actually read them. The model points; it does not conclude.
- Do not treat a null result as a pass in a thin-volume product. Run the file-based analysis instead.
- Keep the model reproducible. If you cannot re-run last year's analysis and get last year's numbers, you cannot show an examiner how a conclusion was reached.
Comply Fair Lending runs logistic regression for underwriting and linear regression for pricing, with field values transformed and dummy-coded rather than by hand, consistency-verification tools for relationships between variables, and application-level residuals that are filterable and usable in a scorecard — which is the mechanism behind the “use residuals to select files” point above. Every view and process saves as a reusable definition, which is what makes a prior year's analysis re-runnable. Definitions for the statistical terms used here are in our compliance glossary.
Sources
- Comptroller’s Handbook, Consumer Compliance: Fair Lending, Version 1.0, Office of the Comptroller of the Currency, January 2023, as amended 14 July 2025. Quotations from page 19 of the booklet. An earlier edition is also online and is stamped RESCINDED on every page; read the current booklet. Current booklet at the OCC.
- Interagency Fair Lending Examination Procedures, Federal Financial Institutions Examination Council, August 2009 (current; replaced the March 1994 procedures). Copy hosted by NCUA (PDF).
Quotations were read from the primary documents on 18 August 2026 and checked against the rendered page image for strikethrough. Examination procedures change; verify against the current edition before relying on any quotation here, including ours.
Frequently Asked Questions
What can a fair lending regression model actually prove?
That a difference in outcome between a prohibited basis group and a control group persists after accounting for the legitimate factors you put into the model, and whether that difference is large enough to be unlikely to have arisen by chance. That is a genuinely useful finding and it is more than a raw approval-rate gap can tell you. What it is not is a finding that the difference is caused by the prohibited basis characteristic. A model measures what remains unexplained by the variables it was given, which is not the same as what is unexplainable.
What can a regression model not prove in a fair lending analysis?
It cannot rule out that a legitimate factor you left out of the model explains the disparity. This is the omitted variable problem and it is the central limitation of the method. The OCC Comptroller's Handbook publishes a three-part test for assessing a model's reliability when a factor is missing or inaccurate: whether the omitted factor is highly correlated with the prohibited basis characteristic, whether it has a strong effect on the outcome, and whether it has a systematic effect. A model also cannot tell you which specific files to remediate, or whether a difference reflects a policy or individual discretion.
Do regulators use the word regression?
Not in the current OCC Comptroller's Handbook on fair lending, where the word regression does not appear at all. The handbook says statistical modeling. The FFIEC Interagency Fair Lending Examination Procedures say statistical analysis, naming it alongside comparative file review as the approaches available to examiners. Regression is the usual implementation of statistical modeling in this context, so the terms describe the same territory, but a vendor feature list that says regression and an examination procedure that says statistical modeling are not using different concepts.
Can a file review overturn a statistical finding of disparity?
Not on its own, and the OCC Comptroller's Handbook is explicit in both directions. It states that the goal of a comparative file review conducted alongside statistical modeling "is not to replicate or replace the statistical model's results," but to find data errors or bank policies incorrectly accounted for in the model, and to identify factors that could be incorporated into it. It then cautions that "it is important not to dismiss statistical results based simply on findings from a manual file review." And a factor you want added to the model "must be legitimately derived from bank policies and occur sufficiently systematically to warrant changes to the model." A handful of individually explicable files is not a rebuttal.
Does statistical significance mean a fair lending violation?
No. Statistical significance is a statement about how unlikely a result is to have arisen by chance, given the model. It is not a legal conclusion, and it does not identify a cause. A statistically significant disparity is a reason to investigate: to check the data, to test whether a legitimate factor was omitted, and to look at the actual files. Treating a significance threshold as the finding itself is the most common way a self-assessment produces either false alarm or false comfort.
This page may be republished with attribution to RATA Associates and a link to rataassociates.com. If you believe a quotation or reading here is wrong, tell us — we would rather fix it than be cited incorrectly.
