The four-fifths rule says that if a group’s selection rate is less than 80% of the rate of the group with the highest rate, enforcement agencies will generally regard that as evidence of adverse impact. It is a screening heuristic, not a legal threshold — smaller differences can still be adverse impact, and larger ones are not automatically unlawful.
- Calculate selection rate per group, then divide each by the highest group’s rate. Below 0.80 is a flag.
- It is explicitly a rule of thumb: statistically significant smaller gaps can count, and small-sample gaps may not.
- Run it at every stage, not just on the final hire. Overall numbers hide the stage doing the damage.
- Passing it is not a defence. A selection procedure still has to be job-related and consistent with business necessity.
- Fixing the ratio by moving a threshold is tuning for the metric. Fix the instrument instead.
This is general information, not legal advice. Adverse-impact analysis is fact-specific and jurisdiction-specific. Engage qualified counsel and, for anything consequential, a statistician.
The arithmetic
Three steps, and you can do it in a spreadsheet.
1. Selection rate per group = selected ÷ applied. 2. Find the highest rate. That group becomes the reference. 3. Impact ratio = each group's rate ÷ the highest rate.
Anything below 0.80 is a flag.
| Group | Applied | Advanced | Selection rate | Impact ratio | |
|---|---|---|---|---|---|
| Group A | 600 | 180 | 30.0% | 1.00 | reference |
| Group B | 450 | 108 | 24.0% | 0.80 | at the line |
| Group C | 380 | 68 | 17.9% | 0.60 | flag |
| Group D | 55 | 8 | 14.5% | 0.48 | flag — but see the count |
Group C is a genuine finding: 380 applicants is a real denominator, and a 0.60 ratio is a substantial gap. Group D is a question, not a finding — with 55 applicants, three additional selections would move the ratio to 0.66 and six would take it past 0.80. Never report a ratio without its counts.
What the rule actually says — and does not
The four-fifths rule comes from the Uniform Guidelines on Employee Selection Procedures, and the Guidelines are careful in a way that most citations of the rule are not. Two qualifications sit right next to it:
- Smaller differences can still be adverse impact where they are statistically significant and practically meaningful. A ratio of 0.85 across 50,000 applicants can be a much stronger signal than 0.55 across 40.
- Larger differences may not be adverse impact where special circumstances apply — small numbers being the main one.
So the rule is a triage heuristic. It is deliberately crude, deliberately conservative, and designed to tell you where to investigate. Treating 0.80 as a pass mark inverts its purpose: you end up managing to the threshold instead of examining the process.
Run it at every stage
This is the mistake that hides most real problems. Teams compute one ratio on final hires, get a comfortable number, and stop.
Adverse impact is usually concentrated in one gate. And stage-level effects compound multiplicatively, so a pipeline where each of four stages has a modest 0.90 ratio ends up around 0.66 overall — from four stages that individually all "passed".
Compute it at every transition:
- Application → screen
- Screen → assessment
- Assessment → interview
- Interview → offer
- Offer → acceptance
That last one is worth its own attention. An adverse pattern in acceptances is not a selection problem — it is a signal about compensation, reputation, or how candidates experienced your process.
The other cut: not just who passed, but who finished
Selection rates only count people who completed a stage. If one group abandons your assessment at twice the rate of another, that group has been filtered — just not by a decision you recorded.
Track completion rates by group alongside selection rates. Common causes of differential drop-off:
- A long unpaid take-home, which selects for who has spare hours
- Assessment requiring specific equipment, bandwidth, or a quiet room
- Rigid scheduling windows that assume a flexible current job
- Invasive proctoring that some candidates decline on principle or for accessibility reasons — see why webcam proctoring is the wrong fix
None of these appear in a selection-rate analysis. All of them shape who you hire.
Diagnosing a flag
You have a 0.60 at the assessment stage. Do not touch the threshold yet.
1. Confirm the data. Is the group assignment right? Are unknowns handled consistently? Is the date range comparable?
2. Localise within the stage. Which items or dimensions drive the gap? Frequently one section is responsible, and often it is the one measuring something adjacent to the job — speed under time pressure, a specific dialect of jargon, or comfort with a test format.
3. Ask what that component measures. Is it job-related? Can you state the connection between the component and performance in the role? If not, you have found something to cut, and cutting it improves validity as well as the ratio.
4. Look for a less-discriminatory alternative that measures the same job-relevant ability. Whether one existed is a question you should be able to answer.
5. Only then consider thresholds — and if you change one, be able to explain the change in terms of what is being measured. "We lowered the cut score until the ratio cleared 0.80" is tuning for the metric, and it will read exactly that way to anyone who examines it.
6. Keep the history. Ratios over time are far more informative than any single snapshot, and the trend is what demonstrates you were paying attention.
Where the data comes from
You cannot compute any of this without demographic data, and collecting it badly creates its own problem.
The safe pattern: collect it through a separate, voluntary self-identification channel, keep it aggregated for analysis, and make it structurally unavailable to anyone making or scoring a decision. Not "policy says don't look" — actually unavailable. Demographic data in the same view as an assessment is an invitation, and the whole point is to measure outcomes rather than let anyone act on the attribute.
If you have no demographic data, that is itself a finding: you are unable to detect adverse impact, which is not a neutral position.
How CalHire is set up for this
Two design properties matter here.
The evaluation surface cannot key on protected attributes, because it never receives them. Identity is stripped at a hard architectural boundary before scoring, the text interview and ranking. Nationality is a specific case worth naming: in Emiratization reporting it may inform sourcing and compliance counting, but it is never a scoring input — that separation is enforced in the product, not left to configuration.
The measurement exists without a data project. The compliance console produces bias-audit exports and stage-level reporting from the same records the hiring decisions came from, rather than a parallel spreadsheet reconstructed later. Independent third-party audit results are published at /bias-audit, and the console is described on the enterprise page.
Anonymous evaluation removes a mechanism of bias — a scorer cannot act on a signal it never received. It does not guarantee equal outcomes, because a genuinely skill-based measure can still correlate with unequal access to skill-building. Which is exactly why you still run the ratios.
Frequently asked questions
- How do you calculate the four-fifths rule?
- For each group, divide the number selected by the number who applied to get a selection rate. Identify the group with the highest rate. Divide every other group’s rate by that highest rate. Any result below 0.80 — that is, below four-fifths — is generally regarded as evidence of adverse impact and warrants investigation.
- Is the four-fifths rule a legal standard?
- No. The Uniform Guidelines present it as a practical rule of thumb that federal enforcement agencies generally apply, and they expressly note that smaller differences may constitute adverse impact where they are statistically significant and practically meaningful, and that larger differences may not where special circumstances apply. Treat it as a screen that tells you where to look.
- What if our sample is small?
- Ratios computed on small numbers swing wildly — with 20 applicants in a group, two people can move the ratio by 0.20. Report the raw counts alongside every ratio, and be sceptical of both alarming and reassuring results at low volume. Aggregate across time or across similar roles to get a usable denominator, and document how you did it.
- We passed the four-fifths test. Are we safe?
- No. Passing means one screening heuristic did not flag your data at that stage in that period. Liability turns on your actual outcomes and on whether your selection procedures are job-related and consistent with business necessity, including whether a less-discriminatory alternative was available. The ratio is a smoke detector, not a certificate.