calhire
Capability

Work-sample assessments

The oldest idea in selection research, done properly: ask someone to do a small version of the job, and score what they produce.

A work-sample assessment asks a candidate to complete a bounded, realistic task from the role and scores the output against a rubric defined in advance. On CalHire it sits alongside the baseline skills assessment for roles where knowledge questions are insufficient, uses the same task and rubric for every candidate, is time-boxed, and is reviewed without the candidate’s identity attached.

Last reviewed

The short version

  • Work samples are among the strongest predictors of job performance in the personnel research literature.
  • Same task, same rubric, same time box for every candidate. Variation is what destroys comparability.
  • Time-boxed on purpose, because an unbounded take-home selects for who has free evenings.
  • Reviewed anonymously, so the output is judged rather than the person who submitted it.

Why work samples, and why they so often go wrong

The case for work samples is unusually strong and unusually old. Decades of meta-analytic research on selection methods repeatedly place work samples near the top for predictive validity, above unstructured interviews and well above years of experience. The reason is not subtle: the closer the assessment resembles the job, the better it predicts the job. Nothing else in selection research is that tidy.

What goes wrong is almost always the implementation. The take-home that quietly expands to twelve hours. The task that is really unpaid work on a live problem. The brief so vague that half the candidates solve a different problem. The rubric written after the submissions arrive, which is another way of saying the rubric was written to justify a preference.

Each of those failures selects for something other than skill. Every one is avoidable. An open-ended task selects for available time, which selects against carers and people already working. A vague brief selects for candidates who share the author’s assumptions. A retrofitted rubric selects for whoever the reviewer already liked.

The rules that make a work sample fair

Bounded, and honest about the bound
A stated time box that reflects the actual work. If the task takes three hours, say three hours and mean it. Candidates who spend nine anyway are a signal about your brief, not about them.
The same task for everyone
Different tasks cannot be compared, however similar they look. One task per role, one rubric, one scale.
A rubric written first
Criteria and weights defined before the first submission. A rubric written afterwards is a rationalisation with a table around it.
Not your real backlog
A task derived from real work, not a live ticket. Asking candidates to do unpaid production work is exploitative, and it also produces worse signal, because the constraints of real work are invisible to someone outside it.
Reviewed without identity
The reviewer sees the submission, not the person. This is the same protection the rest of the platform applies, extended to the one artefact that most invites a reviewer to infer background.

Where it fits in a pipeline

  • After the baseline, not instead of it

    The baseline assessment narrows the field cheaply. A work sample costs both sides real time, so it belongs where the field is already small.

  • Scored into the composite

    The result feeds the ranking alongside test, interview and role fit, on the weights you set for the role.

  • Its own validity window

    Work-sample results carry their own expiry, for the same reason baseline scores do. A dated result is evidence; an undated one is a rumour.

  • Still anonymous

    A work sample does not trigger a reveal. Candidates can demonstrate real capability without identifying themselves to an employer they may not proceed with.

What this does not do

It does not suit every role. Work samples fit best where the output is discrete, gradeable and can be produced alone in a few hours. Roles whose core skill is a live interaction, sustained collaboration, or physical craft are measured badly this way.

It costs candidates real time, and that cost is not distributed evenly. Every additional hour you ask for excludes someone with caring responsibilities or a second job. The time box is a fairness control, not an administrative detail, and stretching it has consequences for who ends up in your pipeline.

It does not defeat outside help by itself. A candidate can get assistance on a take-home. Of course they can. The integrity layer contributes content-authenticity signals, and a follow-up conversation about the submission remains the most reliable check, because explaining a solution requires having understood it.

A work sample cannot rescue a bad rubric. If the criteria reward the approach the author happens to prefer, the assessment will consistently and defensibly select for people who think like the author. Adverse-impact measurement applies to work samples exactly as it does to everything else.

Questions people actually ask

How long should a work sample take?
As short as it can be while still producing signal. Most well-designed samples land between one and three hours. Beyond that you are measuring available time as much as ability, and you will lose good candidates who have none.
Should candidates be paid for work samples?
For a short, synthetic task most organisations do not, and that is a defensible position. If the task is long enough to feel like work, or the output has value to you, payment is the honest answer. That is a policy decision for you rather than a platform setting.
Can candidates use AI tools on a work sample?
Decide and state it in the brief, because both answers are legitimate depending on the role. If the job uses these tools daily, banning them measures something the job does not require. If you do ban them, the content-authenticity signals give you something to review.
Does the reviewer know who produced a submission?
No. Work samples are reviewed anonymously like the rest of the pipeline, so the submission is judged on its own terms.
How does it affect the composite score?
It contributes to the ranking on the weights you set for the role. As with every other component, it informs a decision a person makes rather than making one.

Where this connects to the rest of the platform.

  • Verified skills assessment

    One supervised assessment produces a skills profile an employer can check, instead of a resume they have to take on trust.

  • Composite scoring and blind ranking

    A composite score you set the weights for, defaulting to 40 percent test, 35 percent interview and 25 percent role fit. It orders a list. It never makes a decision.

  • Assessment integrity without surveillance

    Four families of signal, each with a confidence level and PII-free evidence. No webcam, no verdicts, and environmental flags are excluded from scoring where accommodations apply.

Read the reasoning

The evidence and the argument behind what is on this page.

Hiring Playbook8 min read

Does your assessment predict anything? Validity for non-specialists

An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.

Read
Hiring Playbook8 min read

Structured interviews: the highest-return change most teams never make

Same questions, same order, same rubric, scored independently. Structured interviews are unglamorous, uncomfortable to adopt, and the best value in hiring.

Read
Hiring Playbook8 min read

What skills-based hiring actually is (and what it is not)

Skills-based hiring decides on demonstrated ability, not proxies like degrees or job titles. What it means in practice, where it fails, and how to run it well.

Read
Hiring Playbook7 min read

How to write a skills-based job description

Most job descriptions are a wish-list of proxies. How to write one around measurable skills, plus the language patterns that quietly shrink your applicant pool.

Read

Browse all topics on the blog

See a verified pipeline for one of your roles

Post a role free and review anonymous, skill-ranked candidates. No card, no sales call to get started.