calhire
Methodology

How assessments are built and scored

What goes into a score, what the score is used for, and what we have not proven yet.

How a sitting is built

Each sitting’s questions are generated fresh by an AI model from the candidate’s target role and declared skills. The model works inside a fixed structure configured on the platform: a set number of multiple-choice, scenario and written case questions, followed by a short AI text interview.

  • Multiple-choice questions12 questions
  • Scenario questions8 questions
  • Written case questions5 questions
  • AI text interview questions7 questions
  • Text interview7 questions

Total time: about 50 minutes.

There is no item bank and no psychometric equating. Different sittings are kept comparable by the fixed structure, the same rubric and the same weights. We have not yet run statistical equivalence studies between sittings, and we list that below as a limit.

How it is scored

Multiple-choice and scenario questions
Marked right or wrong.
Written case questions
Scored by an AI model against a rubric.
Test score
The average across the test’s questions.
Interview score
The average of four dimensions: communication, content depth, reasoning and professionalism.
Role fit
How much of the candidate’s declared skill set the assessment covered.
Composite score
A weighted blend: test 40%, interview 35% and role fit 25% by default. Employers can configure the weights per role.
Human review
A large spread between the sub-scores flags the result for human review rather than letting one number stand on its own.

Scores inform people; they do not decide. Hiring decisions are made by people.

Integrity and accommodations

During a sitting we record behavioural signals such as focus or tab changes, bursts of pasted text, typing cadence and answer similarity. These signals are advisory: they route a result to a human reviewer and never decide an outcome on their own. A candidate can appeal an integrity flag.

Accommodations are available, including extra time, untimed screen-reader sittings and a low-bandwidth mode.

Adverse-impact monitoring

We compare selection rates across self-disclosed demographic groups using the four-fifths rule. Groups below a minimum size are not reported, because ratios from very small groups are not reliable.

This is a monitoring signal, not a certification. An independent bias audit is planned and has not been completed; its status is published on the bias-audit disclosure.

Post-hire outcomes

Employers are asked for a 90-day hire-quality check-in on the people they hire, so that whether scores predict performance on the job can be studied in future. No validity results have been published yet.

Limits

  • There is no published criterion-validity study yet.
  • Questions are AI-generated for each sitting, so content varies between sittings even within the fixed structure, and no statistical equivalence study has been run.
  • Assessments are available in English only.

Browse the assessment librarySee a sample report cardTrust Center