calhire
All posts
Hiring PlaybookAssessmentSkillsMetrics

Does your assessment predict anything? Validity for non-specialists

An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.

CThe CalHire TeamCalHire8 min read

Reviewed by CalHire Compliance, Compliance & Fairness

A hiring assessment is valid to the extent that its scores relate to actual job performance, and reliable to the extent that it produces consistent scores for the same ability. Difficulty, candidate stress, and how sophisticated a test feels are all unrelated to whether it predicts anything.

  • Reliability is consistency; validity is whether the score means what you think. You need both, in that order.
  • A hard test is not a valid test. Difficulty and predictive power are independent properties.
  • Content validity — does the task resemble the actual work — is the one most teams can improve immediately.
  • Record scores and later performance from day one, or you will never be able to check validity at all.
  • Validity evidence and adverse-impact evidence are different questions. Both need answering.

Two words, and why they are not the same

Reliability is consistency. If the same person with the same ability sits your assessment twice, do they score similarly? If two reviewers score the same submission, do they agree? Unreliable measurement is just noise with a number on it.

Validity is meaning. Does a higher score correspond to doing the job better?

The order matters: an unreliable measure cannot be valid, because it is not measuring anything stable. But a highly reliable measure can be completely invalid. A tape measure is extremely reliable and tells you nothing about whether someone can debug a race condition.

This is where most hiring assessments actually fail. Not on rigour — on relevance.

Difficulty is not validity

The most common mistake in assessment design is treating difficulty as a proxy for quality. A test that makes candidates sweat feels like it is separating the strong from the weak. Often it is separating people who have recently practised that specific puzzle format from people who have not.

Three properties get conflated constantly:

PropertyWhat it meansWhat it tells you about hiring quality
DifficultyHow many candidates score lowNothing on its own
ReliabilityHow consistent scores areA precondition, not a result
ValidityWhether scores relate to job performanceThis is the only one you actually want

A brutal test with a 5% pass rate and no relationship to the work is worse than a moderate test with a real one, because it produces confident wrong decisions and an unnecessarily miserable candidate experience.

The four kinds of evidence you can actually gather

You do not need a psychometrician to make progress. You need to know which type of evidence you are relying on.

1. Content validity — does the task resemble the work? The most practical form, and the one most teams can fix this quarter. Write down the role's actual tasks, then map each assessment item to one. Items that map to nothing get cut. Tasks with no item either do not matter or are unmeasured — decide which.

2. Criterion validity — do scores relate to later performance? The gold standard, and the one almost nobody has, because it requires data you can only collect over time. Which is the point: start recording assessment scores alongside performance outcomes now, so that in eighteen months you can answer the question. Teams that skip this are permanently unable to evaluate their own process.

3. Construct validity — are you measuring the thing you named? A "communication" exercise conducted entirely in writing under time pressure may be measuring typing speed and English fluency. Both may be irrelevant to the role.

4. Face validity — does it look credible to candidates? The weakest evidentially, and it still matters, because it drives completion rates and whether your strongest candidates take you seriously. Just never mistake it for the others.

Four ways assessments quietly break

It tests the format, not the skill. Timed multiple-choice under pressure measures a particular kind of test-taking. If the job is never like that, the score is partly noise.

The rubric is written after the answers arrive. Then it describes the candidate you already liked. Write scoring bands first, including what a mediocre answer looks like.

Scoring drifts. Reviewer standards move over weeks, so a candidate assessed in March is not compared to one assessed in June. Anchor examples and periodic calibration are the fix.

It leaks identity. If the reviewer can see who wrote the submission, the score is measuring the work plus everything they infer from the name, photo, school and employer. This is not a character flaw in your reviewers — it is what happens when irrelevant information is available. Which is why blind evaluation is a measurement decision before it is an ethical one.

Validity is not the same question as fairness

These get merged, and they should not be:

  • Validity: do scores predict performance?
  • Adverse impact: do selection rates differ materially across groups?

You can have the first without the second being acceptable. A measure that predicts performance reasonably well can still screen out one group at a much lower rate, and "but it's predictive" does not resolve that on its own — the question of whether a selection procedure is job-related and consistent with business necessity is exactly what regulators examine. Check both. The mechanics of the second are in the four-fifths rule and adverse impact.

A checklist you can run this week

  1. List the role's real tasks. From someone doing the job, not from the job ad.
  2. Map every assessment item to a task. Cut the orphans.
  3. Check reliability cheaply. Have two reviewers score ten past submissions blind. If they disagree badly, your rubric is the problem, not your reviewers.
  4. Write the rubric first, always. With bands and anchor examples.
  5. Start the score-to-performance log today. Two columns and a date. That is enough to begin.
  6. Run selection rates by group at every stage, not just the final one.
  7. If you use a vendor, ask for the technical documentation — reliability figures, validity evidence, adverse-impact analysis. A vendor who cannot produce it has told you something.

What CalHire does about it

Two design choices in the platform bear directly on validity:

Content validity by construction. The assessment is generated from the skills set for the role, and the composite score is built from a skills test, a text interview, and role-fit — how the verified skills match that specific role. There is deliberately no résumé-match component, because keyword overlap with a self-reported document is not evidence about a person's ability.

The scoring surface sees skills only. Identity is stripped at a hard architectural boundary before any model is involved, across scoring, the text interview, and ranking. Name, photo, school and employer never reach the thing producing the number, so they cannot influence it.

Employers set the composite weights per role — the default is test 40 / interview 35 / role-fit 25 — and below-threshold candidates are flagged for human review rather than auto-rejected, because a threshold is a policy choice and should be made by a person. Details are on the features page.

Validity is not a certificate you obtain once. It is a claim you keep checking. The teams that get it right are simply the ones who wrote the numbers down early enough to check.

Frequently asked questions

What is the difference between reliability and validity?
Reliability is consistency: the same candidate with the same ability should get a similar score on a second sitting, and two reviewers scoring the same work should broadly agree. Validity is meaning: does a higher score correspond to better performance on the job? A test can be highly reliable and completely invalid — it measures something precisely, just not the thing you care about.
Do we need a formal validation study?
For most teams, no — but you do need evidence. Content validity is the practical starting point: document the link between what the role requires and what the assessment asks. If you use a vendor test, ask for their technical documentation, including reliability figures and any adverse-impact analysis. If they cannot produce it, that is your answer.
Is a work sample better than an aptitude test?
A well-built work sample usually has stronger content validity — it resembles the job — and is easier to justify to candidates and regulators. Its weakness is scale: work samples are expensive to score consistently, which is where rubrics and structured scoring matter. Neither format is inherently valid; construction and consistency decide it.
Can an assessment be valid and still be unfair?
Yes, and this is the distinction teams most often miss. Validity asks whether scores predict performance. Adverse impact asks whether selection rates differ across groups. A measure can predict performance reasonably well and still screen out one group at a materially different rate, which is a separate problem requiring separate analysis.
Share this post

Keep reading

Hiring Playbook7 min

Why candidates get ghosted, and how to actually stop

Ghosting is almost never malice. It is what happens when closing the loop is optional, unowned and manual. The four structural causes, and the fix for each.

Hiring Playbook7 min

Dropping the degree requirement: what has to replace it

Removing a degree requirement without replacing the signal makes hiring more subjective, not less. What the degree was doing, and what to measure instead.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.