calhire
All posts
ComplianceComplianceBiasAI in hiring

What an AI bias audit actually measures

A bias audit is an outcome analysis, not a code review. What impact ratios are, what an audit cannot tell you, and how to read a summary you have been handed.

CCalHire ComplianceCompliance & Fairness8 min read

A bias audit measures outcomes, not code. It calculates selection or scoring rates for each demographic category, divides each by the rate of the most-selected group to produce impact ratios, and reports them. It tells you whether a tool selected groups at different rates — not why, and not whether the tool is any good at predicting performance.

  • An audit is arithmetic on outcomes: selection rates, then ratios against the highest-scoring group.
  • It examines sex and race/ethnicity categories, including intersectional combinations.
  • An audit cannot tell you whether a tool is valid, whether it is lawful, or why a disparity exists.
  • Small sample sizes make ratios unstable — read the counts before you read the ratios.
  • Passing an audit is not a defence to a discrimination claim. They answer different questions.

This is general information, not legal advice. Audit requirements and methodologies differ by jurisdiction. Engage a qualified independent auditor and your own counsel.

It is arithmetic on outcomes, not an inspection of the model

The single most common misconception is that a bias audit examines how a system works. It does not. An auditor does not read your model weights, review your training data, or evaluate your architecture.

A bias audit asks one question: did this tool select or score groups of people at different rates?

That is measured by counting. Which is a strength — outcomes are what affect people, and outcome analysis works on systems too complex to introspect — and a limitation, because a number that tells you what happened tells you nothing about why.

The calculation

Two steps.

Step 1 — selection rate per group. For each demographic category, the proportion who were selected.

Step 2 — impact ratio. Each group's rate divided by the rate of the group with the highest rate.

GroupApplicantsSelectedSelection rateImpact ratio
Group A50025050%1.00 (reference)
Group B40016040%0.80
Group C3009030%0.60
Group D40820%0.40

The reference is always the best-performing group, so ratios are at most 1.00. A ratio of 0.60 means that group was selected at 60% of the rate of the most-selected group.

For tools that produce a score rather than a pass/fail, the same arithmetic is applied to the rate of scoring above the median. Same structure, different numerator.

Audits examine sex categories, race/ethnicity categories, and the intersections — for example the combination of sex and race, which is where disparities frequently appear that neither single dimension reveals. A tool can look even on sex and even on race while being materially uneven on a specific combination of the two.

How to read a summary you have been handed

Read in this order. Most people read it backwards.

1. The counts, before the ratios. Look at Group D above: 40 applicants, 8 selected. Move two people and the ratio swings from 0.40 to 0.60. Small denominators produce dramatic, meaningless ratios in both directions. A "passing" ratio on 25 people is not information.

2. The scope. What tool version, what configuration, which roles, what date range? A summary that does not say is not auditable. If you configure your own thresholds and score weights, an audit of the vendor default may not describe your system at all.

3. The data source. Historical data from the employer's own use, or the vendor's test data? Both are permitted in various regimes; they support very different conclusions. Test data tells you about the instrument; your data tells you about your hiring.

4. Excluded categories and unknowns. Candidates who did not self-identify have to be handled somehow. How they were handled can move the answer, and it must be disclosed.

5. The stage measured. A tool measured at final-offer stage can look clean while the screen it feeds is doing the damage — and vice versa. Ask which decision point was analysed.

6. What the auditor declined to conclude. Good audits are careful about their limits. If a summary reads as an endorsement, be suspicious of it.

Four things an audit cannot tell you

Whether the tool works. Adverse impact and predictive validity are independent properties. A tool can have beautiful impact ratios and no relationship to job performance whatsoever — see assessment validity.

Why a disparity exists. The audit reports the gap. Diagnosing it requires digging into which items, features or stages produce it, which is separate analytical work and usually more useful than the audit itself.

Whether you are compliant. Publishing a required audit satisfies a disclosure obligation. It does not resolve whether your selection procedure is job-related and consistent with business necessity — the standard that governs any test or screen.

Whether it will behave the same next quarter. An audit is a snapshot. Retrain the model, change the threshold, shift the applicant mix, and the numbers move.

What to do with a bad result

A failing ratio is diagnostic information. The instinct is to adjust the threshold until the ratio improves, which is the worst available response — it treats the measurement as the problem.

A better sequence:

  1. Localise it. Which stage, which items, which features? Disparity is rarely uniform across a pipeline; usually one gate is responsible.
  2. Ask what that gate is actually measuring. Frequently something adjacent to the job: test-taking speed, bandwidth, equipment quality, fluency in a specific register, availability of uninterrupted time.
  3. Check for a less-discriminatory alternative that measures the same job-relevant ability. This is both good practice and the question you will be asked.
  4. Fix the instrument, not the number. If you cannot explain the change in terms of what is being measured, you have tuned for the audit.
  5. Re-measure, and keep the history. The trend matters more than any single snapshot.

How CalHire approaches it

Our automated employment-decision tooling is independently bias-audited by a third party, and results are published at /bias-audit with the audit date, a summary of results, and the distribution date. Candidates receive notice before an automated tool is used and may request a human-only alternative.

Two structural choices shape what an audit of CalHire is examining:

  • Identity is not on the evaluation surface. Personal information is stripped at a hard architectural boundary before scoring, the text interview, and ranking. The scoring path never receives name, photo, age, school or employer, so it cannot key on them — directly or as a proxy for something it inferred from them.
  • Nothing is auto-rejected. Below-threshold candidates are flagged for human review, and a person makes and is recorded for every person-affecting decision. There is always a human in the record for an auditor to examine.

Employers can export bias-audit data from the compliance console rather than assembling it by hand — see the enterprise page.

An audit is a smoke detector, not a fire-safety certificate. It tells you something is wrong; it does not tell you the house is safe, and it certainly does not put anything out.

Frequently asked questions

What is an impact ratio?
The selection rate for one group divided by the selection rate of the group with the highest rate. If men are selected at 50% and women at 40%, the impact ratio for women is 0.8. For scoring tools the same arithmetic is applied to the rate of scoring above the median rather than to selection.
What does an audit not tell you?
Whether the tool predicts job performance; why a disparity exists; whether the disparity is legally justifiable; and whether the tool behaves the same on your data as on the audited data. It is one measurement of one property, and it is frequently over-read as a clean bill of health.
Can a tool pass a bias audit and still be discriminatory?
Yes. An audit reports ratios at a point in time on a particular dataset. Liability under discrimination law turns on your actual outcomes and on whether your selection procedure is job-related and consistent with business necessity. A favourable audit is evidence, not immunity.
What if we do not have demographic data?
Then you cannot compute impact ratios, which is itself a finding. Where an audit relies on incomplete data, that limitation must be disclosed in the published summary. Collect demographic data through a separate, voluntary, self-identification channel that is kept out of the evaluation surface entirely — never as an assessment input.
Share this post

Keep reading

Compliance7 min

A practical guide to Emiratization-compliant hiring

UAE mainland firms must reach 10% Emirati hires in skilled roles. Here is how to hit the target with verified talent, blind assessment and audit-ready reporting.

Integrity8 min

Candidates are using AI in your assessments. Now what?

AI-text detectors are unreliable and disproportionately flag non-native speakers. What to do instead: assessment design, behavioural signals, and human review.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.