calhire
Capability

Assessment integrity without surveillance

Remote assessment has a real cheating problem. Watching people through their webcam is a bad answer to it, and this page explains what we do instead.

CalHire’s integrity layer detects likely cheating during an assessment using four families of signal: environmental, behavioural, content-authenticity and identity assurance. Each signal is stored with a confidence level and PII-free evidence, and combines into a risk score that routes a session to human review. There is no webcam proctoring, no room scan, and no automated disqualification. A person makes every integrity decision.

Last reviewed

The short version

  • Four signal families: environmental, behavioural, content authenticity, and identity assurance.
  • No webcam feed, no room scan, no eye tracking. Nothing is recorded from a candidate’s home.
  • Environmental signals are capped low, and on a session with documented accommodations they are stored but score zero.
  • Risk scores route to human review. The system never disqualifies anyone on its own.

The problem, stated honestly

Remote assessment cheating is real. It also got dramatically easier. A candidate with a second device and a general-purpose model can produce a plausible answer to most written technical questions in seconds. Any vendor claiming their remote test is uncheatable is either not looking or not telling you.

The usual answer has been to watch harder: webcam feeds, room scans, gaze tracking, keystroke recording, sometimes a live proctor. It works less well than it appears. It also costs more than it looks. Studies of remote proctoring report meaningful false-positive rates, disproportionate flagging of candidates with dark skin, and heavy penalties for anyone whose home is loud, shared or small. Meanwhile the determined cheat uses a second device that the webcam, pointed at a face, cannot see.

So the design question is not how to watch a candidate more closely. It is which observable properties of the work actually distinguish authentic effort from outside help, and how to present those to a human without pretending they are proof.

The four signal families

Every signal is stored with its category, an optional confidence value, and bounded evidence containing no personal data.

Environmental
Focus lost, tab switches, developer tools opened, browser extensions present, virtual-machine heuristics, second screen, screen mirroring, fullscreen exited. These are the weakest signals and are weighted accordingly, because a tab switch is as likely to be a delivery notification as a search for the answer.
Behavioural
Paste bursts, large clipboard inserts, response latency anomalies, typing cadence anomalies. An answer that arrives fully formed in one paste is worth a look. It is not, on its own, an accusation.
Content authenticity
The signals that actually carry weight. Machine-generated likeness in the answer, similarity across candidates in the same cohort, style inconsistency between sections, failure to handle a clarifying question, and reasoning that contradicts itself. Outside help shows up in the work far more reliably than in the room.
Identity assurance
Liveness check failure and identity mismatch, used where a tenant has enabled identity verification. This answers a different question from the others: not whether help was used, but whether the person who sat the test is the person who applied.

The fairness exclusion, and why it exists

Environmental signals punish circumstances rather than conduct. That is the whole problem with them. A candidate using a screen reader triggers focus events. Someone with a magnifier runs a second window. A candidate with a motor impairment may use dictation software that a naive keystroke analysis reads as anomalous. A candidate sitting an assessment on a borrowed laptop in a shared room accumulates flags for existing.

So a session with documented accommodations has its environmental signals marked as excluded from scoring. The signal is still stored, visible and auditable, because deleting it would be its own kind of dishonesty. It simply contributes zero to the risk score.

The reason this is built in rather than left to reviewer discretion is that reviewer discretion has a track record here. Asking a person to mentally discount a flag they can see, in the middle of comparing candidates, is asking for the failure mode that accommodation policies exist to prevent.

What happens when a session is flagged

  1. 1

    Signals accumulate with evidence attached

    Each one records its type, category, confidence where applicable, and a bounded evidence payload with no personal data in it. That payload is shown verbatim in the review panel rather than summarised into an adjective.

  2. 2

    A risk score is computed

    Weighted by family, with environmental capped low and content-authenticity carrying the most. Accommodation-excluded signals contribute nothing.

  3. 3

    The session routes to human review

    The candidate stays in the pipeline. A reviewer sees the score, every contributing signal and the evidence for each, and can look at the answers directly.

  4. 4

    A person decides, and it is recorded

    Invalidate, re-test, or accept. Whichever it is, it is attributed to a named reviewer in the audit trail, and the candidate is told what happened and why.

Design choices that reduce cheating before detection has to

Detection is the last line. Most of the benefit comes from making cheating less useful in the first place.

  • Generated per candidate

    Items are produced against the candidate’s declared skills rather than pulled from a fixed bank. There is no shared answer key to circulate, because there is no shared paper.

  • Clarifying questions

    A follow-up that probes the reasoning behind an answer is cheap to ask and awkward to fake, because it requires understanding the answer rather than possessing it.

  • Work samples

    For roles where it fits, a realistic task graded on process and outcome. Harder to outsource convincingly, and better evidence when it goes well.

  • The score expires

    Ninety-day validity limits the value of a single fraudulent result. A cheated score is a wasting asset rather than a permanent credential.

What this does not do

It does not catch every cheat. A patient candidate with a second device, working slowly and rewriting in their own words, can defeat all four families. What the layer changes is the effort required, which is enough to shift most casual cheating without punishing everyone else to catch the determined few.

It does not produce proof. Every signal is probabilistic, machine-text detection especially so. A high risk score means look closely, never conclude. Treating it as a verdict would reproduce exactly the harm that proctoring software is criticised for.

It cannot confirm identity by itself. Behavioural and content signals say something about the work, not about who produced it. Where identity actually matters, the identity-assurance flow is a separate, deliberate check available to enterprise tenants.

It creates a reviewer workload. Flags need reading, and a team that ignores them is running detection theatre. If nobody has time to review, turn the thresholds down rather than accumulating unread flags, which is the worse of the two failures.

Questions people actually ask

Do you record the candidate’s webcam or screen?
No. There is no video capture, no screen recording and no room scan anywhere in the assessment flow. Environmental signals are browser events such as focus changes, not footage.
Can a candidate be disqualified automatically for a high risk score?
No. Every integrity outcome is decided by a person, recorded against their name, and communicated to the candidate. This holds regardless of how high the score is.
What about candidates who need accommodations?
A session with documented accommodations has its environmental signals excluded from scoring. They are still stored and visible for transparency, but they contribute zero to the risk score.
How accurate is AI-text detection?
Imperfect, in both directions, and the platform treats it accordingly. Machine-likeness is one signal among several, always accompanied by its confidence level, and never sufficient on its own for a decision.
Does the candidate know what is being monitored?
Yes. Candidates are told before they begin what the integrity layer observes and what it does not. Detection that depends on candidates not knowing about it is not detection, it is a trap.
Can we see the evidence behind a flag?
Yes. The bounded evidence for each signal is shown verbatim in the review panel, with no personal data in it, so a reviewer weighs the actual observation rather than a label.

Where this connects to the rest of the platform.

  • Verified skills assessment

    One supervised assessment produces a skills profile an employer can check, instead of a resume they have to take on trust.

  • AI text interview

    A structured, text-only interview that asks every candidate the same core questions and scores against a fixed rubric. No video, and no personal data ever reaches the model.

  • AI governance, bias auditing and the audit trail

    Adverse-impact analysis on the four-fifths rule, a hash-chained audit trail, candidate appeals, and DSAR handling. The evidence exists before anyone asks for it.

Read the reasoning

The evidence and the argument behind what is on this page.

Integrity8 min read

Candidates are using AI in your assessments. Now what?

AI-text detectors are unreliable and disproportionately flag non-native speakers. What to do instead: assessment design, behavioural signals, and human review.

Read
Integrity8 min read

Why webcam proctoring is the wrong fix for assessment integrity

Webcam proctoring flags disabled candidates, people with darker skin and anyone without a private room. What it costs, and what to measure instead.

Read
Integrity8 min read

Fake candidates: proxy interviews, stolen identities and deepfakes

Remote hiring created a real identity-fraud problem: proxy interviewers, borrowed identities and deepfaked video. How the fraud works and where to verify.

Read
Hiring Playbook8 min read

Does your assessment predict anything? Validity for non-specialists

An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.

Read

Browse all topics on the blog

See a verified pipeline for one of your roles

Post a role free and review anonymous, skill-ranked candidates. No card, no sales call to get started.