Work-sample assessments
The oldest idea in selection research, done properly: ask someone to do a small version of the job, and score what they produce.
A work-sample assessment asks a candidate to complete a bounded, realistic task from the role and scores the output against a rubric defined in advance. On CalHire it sits alongside the baseline skills assessment for roles where knowledge questions are insufficient, uses the same task and rubric for every candidate, is time-boxed, and is reviewed without the candidate’s identity attached.
Last reviewed
The short version
- Work samples are among the strongest predictors of job performance in the personnel research literature.
- Same task, same rubric, same time box for every candidate. Variation is what destroys comparability.
- Time-boxed on purpose, because an unbounded take-home selects for who has free evenings.
- Reviewed anonymously, so the output is judged rather than the person who submitted it.
Why work samples, and why they so often go wrong
The case for work samples is unusually strong and unusually old. Decades of meta-analytic research on selection methods repeatedly place work samples near the top for predictive validity, above unstructured interviews and well above years of experience. The reason is not subtle: the closer the assessment resembles the job, the better it predicts the job. Nothing else in selection research is that tidy.
What goes wrong is almost always the implementation. The take-home that quietly expands to twelve hours. The task that is really unpaid work on a live problem. The brief so vague that half the candidates solve a different problem. The rubric written after the submissions arrive, which is another way of saying the rubric was written to justify a preference.
Each of those failures selects for something other than skill. Every one is avoidable. An open-ended task selects for available time, which selects against carers and people already working. A vague brief selects for candidates who share the author’s assumptions. A retrofitted rubric selects for whoever the reviewer already liked.
The rules that make a work sample fair
- Bounded, and honest about the bound
- A stated time box that reflects the actual work. If the task takes three hours, say three hours and mean it. Candidates who spend nine anyway are a signal about your brief, not about them.
- The same task for everyone
- Different tasks cannot be compared, however similar they look. One task per role, one rubric, one scale.
- A rubric written first
- Criteria and weights defined before the first submission. A rubric written afterwards is a rationalisation with a table around it.
- Not your real backlog
- A task derived from real work, not a live ticket. Asking candidates to do unpaid production work is exploitative, and it also produces worse signal, because the constraints of real work are invisible to someone outside it.
- Reviewed without identity
- The reviewer sees the submission, not the person. This is the same protection the rest of the platform applies, extended to the one artefact that most invites a reviewer to infer background.
Where it fits in a pipeline
After the baseline, not instead of it
The baseline assessment narrows the field cheaply. A work sample costs both sides real time, so it belongs where the field is already small.
Scored into the composite
The result feeds the ranking alongside test, interview and role fit, on the weights you set for the role.
Its own validity window
Work-sample results carry their own expiry, for the same reason baseline scores do. A dated result is evidence; an undated one is a rumour.
Still anonymous
A work sample does not trigger a reveal. Candidates can demonstrate real capability without identifying themselves to an employer they may not proceed with.
What this does not do
It does not suit every role. Work samples fit best where the output is discrete, gradeable and can be produced alone in a few hours. Roles whose core skill is a live interaction, sustained collaboration, or physical craft are measured badly this way.
It costs candidates real time, and that cost is not distributed evenly. Every additional hour you ask for excludes someone with caring responsibilities or a second job. The time box is a fairness control, not an administrative detail, and stretching it has consequences for who ends up in your pipeline.
It does not defeat outside help by itself. A candidate can get assistance on a take-home. Of course they can. The integrity layer contributes content-authenticity signals, and a follow-up conversation about the submission remains the most reliable check, because explaining a solution requires having understood it.
A work sample cannot rescue a bad rubric. If the criteria reward the approach the author happens to prefer, the assessment will consistently and defensibly select for people who think like the author. Adverse-impact measurement applies to work samples exactly as it does to everything else.
Questions people actually ask
How long should a work sample take?
Should candidates be paid for work samples?
Can candidates use AI tools on a work sample?
Does the reviewer know who produced a submission?
How does it affect the composite score?
Related
Where this connects to the rest of the platform.
Verified skills assessment
One supervised assessment produces a skills profile an employer can check, instead of a resume they have to take on trust.
Composite scoring and blind ranking
A composite score you set the weights for, defaulting to 40 percent test, 35 percent interview and 25 percent role fit. It orders a list. It never makes a decision.
Assessment integrity without surveillance
Four families of signal, each with a confidence level and PII-free evidence. No webcam, no verdicts, and environmental flags are excluded from scoring where accommodations apply.
Read the reasoning
The evidence and the argument behind what is on this page.
Does your assessment predict anything? Validity for non-specialists
An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.
ReadStructured interviews: the highest-return change most teams never make
Same questions, same order, same rubric, scored independently. Structured interviews are unglamorous, uncomfortable to adopt, and the best value in hiring.
ReadWhat skills-based hiring actually is (and what it is not)
Skills-based hiring decides on demonstrated ability, not proxies like degrees or job titles. What it means in practice, where it fails, and how to run it well.
ReadHow to write a skills-based job description
Most job descriptions are a wish-list of proxies. How to write one around measurable skills, plus the language patterns that quietly shrink your applicant pool.
ReadSee a verified pipeline for one of your roles
Post a role free and review anonymous, skill-ranked candidates. No card, no sales call to get started.