BISG Proxy Methodology: How It Works and Where It Fails
Published 18 August 2026 · Quoted from “Using publicly available information to proxy for unidentified race and ethnicity: A methodology and assessment,” Consumer Financial Protection Bureau, Summer 2014 — the paper that made BISG the de facto standard. Linked at the foot of the page.
The short answer. BISG combines a surname probability from the Census Bureau surname list with the racial and ethnic composition of the applicant’s geography, using Bayes’ theorem. It returns a probability for each category, not an assignment to one.
The most consequential mistake in using it is forgetting that. In the CFPB’s own test pool, summing the probabilities overestimated Hispanic applicants by 4%. Applying an 80% threshold to label people instead underestimated them by 25%. The method is more accurate when you resist the urge to turn it into labels.
How it works
Two independent pieces of public information, combined:
- Surname. The Census Bureau surname list gives, for a given surname, the distribution across race and ethnicity categories. Surnames are first standardised — special characters and titles such as JR and SR removed, compound names parsed.
- Geography. Census data gives the racial and ethnic composition of the adult population (18 and over) in the area where the applicant lives.
Bayes’ theorem updates the surname distribution using the geographic distribution, producing a probability for each category that sums to one across categories. RATA also derives an ethnicity proxy from surname and a sex proxy from forename, which are separate estimates built on the same principle.
Geocoding precision sets the ceiling
This is the part most often missed, and the CFPB states it plainly: “the precision of geocoding determines the precision of the demographic information relied upon.” The hierarchy in the paper is specific:
| How precisely the address geocodes | Demographic data the proxy then uses |
|---|---|
| Exact street address (“rooftop”) | Census block group containing the address — or the census tract, if that block group has zero population |
| Street name, 9-digit ZIP, or 5-digit ZIP | 5-digit ZIP code demographics — substantially coarser |
| Cannot be geocoded, or only to city or state | Excluded from the analysis entirely |
Two consequences worth drawing out. First, a block group is a much smaller and more homogeneous unit than a 5-digit ZIP, so an address that only resolves to ZIP level produces a materially blunter proxy for the same applicant. Second, addresses that fail to geocode do not produce a bad proxy — they produce no proxy, and drop out of the population being analysed. A fair lending analysis with a poor geocoding match rate is therefore not merely less precise; it is quietly analysing a different, smaller population than the one you filed.
That is the honest connection between geocoding and fair lending work, and it is the CFPB’s point rather than ours. How we measure geocoding accuracy, and the difference between match rate and accuracy, is covered separately.
Do not threshold the probabilities
The single most actionable finding in the CFPB paper. Working with a mortgage pool where actual reported race and ethnicity was known, the Bureau compared two ways of using BISG output against the 11,073 applicants who reported as Hispanic:
| Method | Estimate | Error vs reported |
|---|---|---|
| Sum the probabilities across applicants | 11,516 | Overestimates by 4% |
| Apply an 80% threshold to classify each applicant | 8,314 | Underestimates by 25% |
The paper explains why: the underestimation “is driven by the failure to count the large number of individuals… who are reported as being Hispanic in the mortgage sample but for whom the BISG probability of assignment is less than 80%.” And thresholding does not only lose people — it invents them. The paper notes 881 applicants reported as non-Hispanic White who were nonetheless assigned a Hispanic probability of 80% or higher.
So a threshold rule fails in both directions at once, and by a much larger margin than the probabilistic approach. If your analysis converts BISG output into a categorical field because that is easier to group and filter, you have introduced a larger error than the proxy itself carries.
The known directional bias
The CFPB is candid that all three proxies it assessed share a bias:
The magnitude of that mismatch is worth seeing, and the paper supplies both sides. Per the 2010 Census, the US adult population was 14% Hispanic, 67% non-Hispanic White, 12% non-Hispanic Black, 5% Asian/Pacific Islander and 1% American Indian/Alaska Native. Per 2010 HMDA data for all reporting originators, mortgage applicants were 7% Hispanic and 80% non-Hispanic White.
The proxy is built from the population at large and applied to a population that applies for mortgages. Those are different groups, and the difference shows up as directional bias rather than random noise — which matters, because a systematic bias does not average out as your sample grows.
Better than the alternatives, not a solved problem
The paper does find BISG superior to the traditional single-input approaches: the BISG proxy “comes closer to approximating the reported race and ethnicity than the traditional proxy methodologies, with the only exception being for Asian/Pacific Islanders and Multiracial.” But it describes the improvement as “small absolute gains in accuracy… for some groups relative to the traditional methods.” That is the correct register for talking about BISG: the best available tool for a genuinely hard problem, not a substitute for reported data.
If you are using BISG in a self-assessment
- Keep the probabilities as probabilities. Sum them for group-level figures. Resist creating a categorical field, however much easier it makes grouping.
- Know your geocoding match rate at block-group level, not just overall. The share of applicants resolved only to 5-digit ZIP is the share getting a blunt proxy, and the share not geocoded at all has silently left your analysis.
- Expect the directional bias and do not treat a small overrepresentation of a protected group in your proxied data as a data error to be corrected. The CFPB found the same thing.
- Use reported data wherever you have it. The proxy exists for the applicants and products where race and ethnicity were not collected, not as a replacement for those where they were.
- Document which version of the surname list and which census vintage you used. A proxy is not reproducible without them, and a result you cannot reproduce is a result you cannot explain.
Comply Fair Lending includes BISG using the CFPB-approved surname and geocoding approach, and RATA’s BISG proxy service derives the race proxy from a GeoPlus geocode — which is the block-group-precision path in the table above rather than the ZIP-level fallback. Definitions for the terms used here are in our compliance glossary, which distinguishes geocoding in the compliance sense from geocoding as a BISG input, because they are not the same thing.
Source
- Using publicly available information to proxy for unidentified race and ethnicity: A methodology and assessment, Consumer Financial Protection Bureau, Summer 2014. PDF at the CFPB. All figures and quotations above are from this document.
Read from the primary document on 18 August 2026. The paper dates from 2014 and census vintages and surname lists have been updated since; verify current inputs before relying on any figure here, including ours.
Frequently Asked Questions
How does BISG actually work?
BISG starts from two independent pieces of public information. The Census Bureau surname list gives the probability distribution across race and ethnicity categories for a given surname. Census demographic data gives the racial and ethnic composition of the adult population in the applicant's geography. Bayes' theorem combines them: the surname distribution is updated using the geographic distribution to produce a single probability for each category. The output is a set of probabilities per applicant summing to one, not an assignment to a category.
Why does geocoding precision matter for a BISG proxy?
Because it determines which demographic data the proxy can use. The CFPB methodology paper states that the precision of geocoding determines the precision of the demographic information relied upon. An address geocoded to an exact street address uses the census block group containing it, falling back to the census tract if that block group has zero population. An address geocoded only to street name, 9-digit ZIP or 5-digit ZIP uses 5-digit ZIP demographics, which are far coarser. And addresses that cannot be geocoded, or can only be resolved to something less precise than a 5-digit ZIP such as a city or state, are excluded from the analysis entirely.
Should I convert BISG probabilities into race categories using a threshold?
The CFPB's own assessment argues against it, with figures. In its mortgage test pool, summing the BISG probabilities overestimated the number of Hispanic applicants by 4 percent against the reported figure. Applying an 80 percent threshold to classify individuals instead underestimated the reported number by 25 percent, because it discards the large number of applicants who genuinely were Hispanic but whose BISG probability fell below 80 percent. The threshold approach also produced false positives: applicants reported as non-Hispanic White who were nonetheless assigned a Hispanic probability of 80 percent or more. Keeping the probabilities as probabilities was substantially more accurate than forcing classification.
Is BISG biased in a known direction?
Yes, and the CFPB documents it. All three proxies it assessed, including BISG, tend to underestimate the non-Hispanic White population and overestimate the other race and ethnicity categories. The paper attributes this to the proxy being constructed from the racial and ethnic composition of census-based populations while being applied to people applying for mortgages, which is a differently composed group. It gives both distributions: in the 2010 Census, 14 percent of the US adult population was Hispanic and 67 percent non-Hispanic White, whereas in 2010 HMDA data only 7 percent of mortgage applicants were Hispanic and 80 percent non-Hispanic White.
Is BISG better than using surname or geography alone?
Generally yes, per the CFPB's assessment: the BISG proxy comes closer to approximating reported race and ethnicity than the traditional proxy methodologies, with the exception of Asian and Pacific Islander and Multiracial categories. The paper is measured about the size of the improvement, describing small absolute gains in accuracy for some groups relative to the traditional methods. So BISG is the better tool rather than a solved problem.
This page may be republished with attribution to RATA Associates and a link to rataassociates.com. If you believe a figure or reading here is wrong, tell us — we would rather fix it than be cited incorrectly.
