calhire
ComplianceComplianceBiasAI in hiring

What an AI bias audit actually measures

A bias audit is an outcome analysis, not a code review. What impact ratios are, what an audit cannot tell you, and how to read a summary you have been handed.

CCompliance & Fairness8 min read

A bias audit measures outcomes, not code. It calculates selection or scoring rates for each demographic category, divides each by the rate of the most-selected group to produce impact ratios, and reports them. It tells you whether a tool selected groups at different rates, not why, and not whether the tool is any good at predicting performance.

  • An audit is arithmetic on outcomes: selection rates, then ratios against the highest-scoring group.
  • It examines sex and race/ethnicity categories, including intersectional combinations.
  • An audit cannot tell you whether a tool is valid, whether it is lawful, or why a disparity exists.
  • Small sample sizes make ratios unstable, read the counts before you read the ratios.
  • Passing an audit is not a defence to a discrimination claim. They answer different questions.

This is general information, not legal advice. Audit requirements and methodologies differ by jurisdiction. Engage a qualified independent auditor and your own counsel.

It is arithmetic on outcomes, not an inspection of the model

The single most common misconception is that a bias audit examines how a system works. It does not. An auditor does not read your model weights, review your training data, or evaluate your architecture.

A bias audit asks one question: did this tool select or score groups of people at different rates?

That is measured by counting. Which is a strength, outcomes are what affect people, and outcome analysis works on systems too complex to introspect, and a limitation, because a number that tells you what happened tells you nothing about why.

The calculation

Two steps.

Step 1, selection rate per group. For each demographic category, the proportion who were selected.

Step 2, impact ratio. Each group's rate divided by the rate of the group with the highest rate.

GroupApplicantsSelectedSelection rateImpact ratio
Group A50025050%1.00 (reference)
Group B40016040%0.80
Group C3009030%0.60
Group D40820%0.40

The reference is always the best-performing group, so ratios are at most 1.00. A ratio of 0.60 means that group was selected at 60% of the rate of the most-selected group.

For tools that produce a score rather than a pass/fail, the same arithmetic is applied to the rate of scoring above the median. Same structure, different numerator.

Audits examine sex categories, race/ethnicity categories, and the intersections, for example the combination of sex and race, which is where disparities frequently appear that neither single dimension reveals. A tool can look even on sex and even on race while being materially uneven on a specific combination of the two.

How to read a summary you have been handed

Read in this order. Most people read it backwards.

1. The counts, before the ratios. Look at Group D above: 40 applicants, 8 selected. Move two people and the ratio swings from 0.40 to 0.60. Small denominators produce dramatic, meaningless ratios in both directions. A "passing" ratio on 25 people is not information.

2. The scope. What tool version, what configuration, which roles, what date range? A summary that does not say is not auditable. If you configure your own thresholds and score weights, an audit of the vendor default may not describe your system at all.

3. The data source. Historical data from the employer's own use, or the vendor's test data? Both are permitted in various regimes; they support very different conclusions. Test data tells you about the instrument; your data tells you about your hiring.

4. Excluded categories and unknowns. Candidates who did not self-identify have to be handled somehow. How they were handled can move the answer, and it must be disclosed.

5. The stage measured. A tool measured at final-offer stage can look clean while the screen it feeds is doing the damage, and vice versa. Ask which decision point was analysed.

6. What the auditor declined to conclude. Good audits are careful about their limits. If a summary reads as an endorsement, be suspicious of it.

Four things an audit cannot tell you

Whether the tool works. Adverse impact and predictive validity are independent properties. A tool can have beautiful impact ratios and no relationship to job performance whatsoever, see assessment validity.

Why a disparity exists. The audit reports the gap. Diagnosing it requires digging into which items, features or stages produce it, which is separate analytical work and usually more useful than the audit itself.

Whether you are compliant. Publishing a required audit satisfies a disclosure obligation. It does not resolve whether your selection procedure is job-related and consistent with business necessity, the standard that governs any test or screen.

Whether it will behave the same next quarter. An audit is a snapshot. Retrain the model, change the threshold, shift the applicant mix, and the numbers move.

What to do with a bad result

A failing ratio is diagnostic information. The instinct is to adjust the threshold until the ratio improves, which is the worst available response. It treats the measurement as the problem.

A better sequence:

  1. Localise it. Which stage, which items, which features? Disparity is rarely uniform across a pipeline; usually one gate is responsible.
  2. Ask what that gate is actually measuring. Frequently something adjacent to the job: test-taking speed, bandwidth, equipment quality, fluency in a specific register, availability of uninterrupted time.
  3. Check for a less-discriminatory alternative that measures the same job-relevant ability. This is both good practice and the question you will be asked.
  4. Fix the instrument, not the number. If you cannot explain the change in terms of what is being measured, you have tuned for the audit.
  5. Re-measure, and keep the history. The trend matters more than any single snapshot.

What to fix before your next audit

Most of the pain in a bias audit is not the analysis. It is discovering, three weeks in, that the data needed to run it was never captured in a usable form.

Two things are worth putting in place well before an auditor arrives. The first is configuration history, so you can state what the tool was doing on any given date; without it, an audit describes a system that may no longer exist. The second is a decision record: who decided, on what basis, and what they saw. Where every below-threshold candidate is flagged for human review rather than removed automatically, as on CalHire, that record exists by default and an auditor has something to examine. Where a threshold silently eliminates people, there is nothing to look at, which is itself the finding.

Neither is glamorous work and both take a quarter to backfill. Starting them the week before an audit is how a routine exercise becomes an expensive one.

An audit is a smoke detector

It tells you something is wrong. It does not tell you the house is safe, it does not tell you what is burning, and it certainly does not put anything out.

Treat a clean result as the beginning of the diagnostic work rather than the end of it, and treat a failing ratio as information about your process rather than a problem with the measurement. Teams that get this the wrong way round end up tuning numbers, and the tuning is always visible to whoever looks next.

Frequently asked questions

What is an impact ratio?
The selection rate for one group divided by the selection rate of the group with the highest rate. If men are selected at 50% and women at 40%, the impact ratio for women is 0.8. For scoring tools the same arithmetic is applied to the rate of scoring above the median rather than to selection.
What does an audit not tell you?
Whether the tool predicts job performance; why a disparity exists; whether the disparity is legally justifiable; and whether the tool behaves the same on your data as on the audited data. It is one measurement of one property, and it is frequently over-read as a clean bill of health.
Can a tool pass a bias audit and still be discriminatory?
Yes. An audit reports ratios at a point in time on a particular dataset. Liability under discrimination law turns on your actual outcomes and on whether your selection procedure is job-related and consistent with business necessity. A favourable audit is evidence, not immunity.
What if we do not have demographic data?
Then you cannot compute impact ratios, which is itself a finding. Where an audit relies on incomplete data, that limitation must be disclosed in the published summary. Collect demographic data through a separate, voluntary, self-identification channel that is kept out of the evaluation surface entirely, never as an assessment input.
Share this post

Keep reading

Compliance9 min

The EU AI Act and hiring: what recruitment teams are responsible for

AI used to recruit or evaluate candidates is high-risk under the EU AI Act. What that means for employers, what providers owe you, and what to ask vendors.

Compliance9 min

GDPR and candidate data: what recruiting teams get wrong

Consent is usually the wrong lawful basis for recruitment. What to rely on instead, how long to keep applications, and the rules on automated decisions.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.