Assessment integrity without surveillance
Remote assessment has a real cheating problem. Watching people through their webcam is a bad answer to it, and this page explains what we do instead.
CalHire’s integrity layer detects likely cheating during an assessment using four families of signal: environmental, behavioural, content-authenticity and identity assurance. Each signal is stored with a confidence level and PII-free evidence, and combines into a risk score that routes a session to human review. There is no webcam proctoring, no room scan, and no automated disqualification. A person makes every integrity decision.
Last reviewed
The short version
- Four signal families: environmental, behavioural, content authenticity, and identity assurance.
- No webcam feed, no room scan, no eye tracking. Nothing is recorded from a candidate’s home.
- Environmental signals are capped low, and on a session with documented accommodations they are stored but score zero.
- Risk scores route to human review. The system never disqualifies anyone on its own.
The problem, stated honestly
Remote assessment cheating is real. It also got dramatically easier. A candidate with a second device and a general-purpose model can produce a plausible answer to most written technical questions in seconds. Any vendor claiming their remote test is uncheatable is either not looking or not telling you.
The usual answer has been to watch harder: webcam feeds, room scans, gaze tracking, keystroke recording, sometimes a live proctor. It works less well than it appears. It also costs more than it looks. Studies of remote proctoring report meaningful false-positive rates, disproportionate flagging of candidates with dark skin, and heavy penalties for anyone whose home is loud, shared or small. Meanwhile the determined cheat uses a second device that the webcam, pointed at a face, cannot see.
So the design question is not how to watch a candidate more closely. It is which observable properties of the work actually distinguish authentic effort from outside help, and how to present those to a human without pretending they are proof.
The four signal families
Every signal is stored with its category, an optional confidence value, and bounded evidence containing no personal data.
- Environmental
- Focus lost, tab switches, developer tools opened, browser extensions present, virtual-machine heuristics, second screen, screen mirroring, fullscreen exited. These are the weakest signals and are weighted accordingly, because a tab switch is as likely to be a delivery notification as a search for the answer.
- Behavioural
- Paste bursts, large clipboard inserts, response latency anomalies, typing cadence anomalies. An answer that arrives fully formed in one paste is worth a look. It is not, on its own, an accusation.
- Content authenticity
- The signals that actually carry weight. Machine-generated likeness in the answer, similarity across candidates in the same cohort, style inconsistency between sections, failure to handle a clarifying question, and reasoning that contradicts itself. Outside help shows up in the work far more reliably than in the room.
- Identity assurance
- Liveness check failure and identity mismatch, used where a tenant has enabled identity verification. This answers a different question from the others: not whether help was used, but whether the person who sat the test is the person who applied.
The fairness exclusion, and why it exists
Environmental signals punish circumstances rather than conduct. That is the whole problem with them. A candidate using a screen reader triggers focus events. Someone with a magnifier runs a second window. A candidate with a motor impairment may use dictation software that a naive keystroke analysis reads as anomalous. A candidate sitting an assessment on a borrowed laptop in a shared room accumulates flags for existing.
So a session with documented accommodations has its environmental signals marked as excluded from scoring. The signal is still stored, visible and auditable, because deleting it would be its own kind of dishonesty. It simply contributes zero to the risk score.
The reason this is built in rather than left to reviewer discretion is that reviewer discretion has a track record here. Asking a person to mentally discount a flag they can see, in the middle of comparing candidates, is asking for the failure mode that accommodation policies exist to prevent.
What happens when a session is flagged
- 1
Signals accumulate with evidence attached
Each one records its type, category, confidence where applicable, and a bounded evidence payload with no personal data in it. That payload is shown verbatim in the review panel rather than summarised into an adjective.
- 2
A risk score is computed
Weighted by family, with environmental capped low and content-authenticity carrying the most. Accommodation-excluded signals contribute nothing.
- 3
The session routes to human review
The candidate stays in the pipeline. A reviewer sees the score, every contributing signal and the evidence for each, and can look at the answers directly.
- 4
A person decides, and it is recorded
Invalidate, re-test, or accept. Whichever it is, it is attributed to a named reviewer in the audit trail, and the candidate is told what happened and why.
Design choices that reduce cheating before detection has to
Detection is the last line. Most of the benefit comes from making cheating less useful in the first place.
Generated per candidate
Items are produced against the candidate’s declared skills rather than pulled from a fixed bank. There is no shared answer key to circulate, because there is no shared paper.
Clarifying questions
A follow-up that probes the reasoning behind an answer is cheap to ask and awkward to fake, because it requires understanding the answer rather than possessing it.
Work samples
For roles where it fits, a realistic task graded on process and outcome. Harder to outsource convincingly, and better evidence when it goes well.
The score expires
Ninety-day validity limits the value of a single fraudulent result. A cheated score is a wasting asset rather than a permanent credential.
What this does not do
It does not catch every cheat. A patient candidate with a second device, working slowly and rewriting in their own words, can defeat all four families. What the layer changes is the effort required, which is enough to shift most casual cheating without punishing everyone else to catch the determined few.
It does not produce proof. Every signal is probabilistic, machine-text detection especially so. A high risk score means look closely, never conclude. Treating it as a verdict would reproduce exactly the harm that proctoring software is criticised for.
It cannot confirm identity by itself. Behavioural and content signals say something about the work, not about who produced it. Where identity actually matters, the identity-assurance flow is a separate, deliberate check available to enterprise tenants.
It creates a reviewer workload. Flags need reading, and a team that ignores them is running detection theatre. If nobody has time to review, turn the thresholds down rather than accumulating unread flags, which is the worse of the two failures.
Questions people actually ask
Do you record the candidate’s webcam or screen?
Can a candidate be disqualified automatically for a high risk score?
What about candidates who need accommodations?
How accurate is AI-text detection?
Does the candidate know what is being monitored?
Can we see the evidence behind a flag?
Related
Where this connects to the rest of the platform.
Verified skills assessment
One supervised assessment produces a skills profile an employer can check, instead of a resume they have to take on trust.
AI text interview
A structured, text-only interview that asks every candidate the same core questions and scores against a fixed rubric. No video, and no personal data ever reaches the model.
AI governance, bias auditing and the audit trail
Adverse-impact analysis on the four-fifths rule, a hash-chained audit trail, candidate appeals, and DSAR handling. The evidence exists before anyone asks for it.
Read the reasoning
The evidence and the argument behind what is on this page.
Candidates are using AI in your assessments. Now what?
AI-text detectors are unreliable and disproportionately flag non-native speakers. What to do instead: assessment design, behavioural signals, and human review.
ReadWhy webcam proctoring is the wrong fix for assessment integrity
Webcam proctoring flags disabled candidates, people with darker skin and anyone without a private room. What it costs, and what to measure instead.
ReadFake candidates: proxy interviews, stolen identities and deepfakes
Remote hiring created a real identity-fraud problem: proxy interviewers, borrowed identities and deepfaked video. How the fraud works and where to verify.
ReadDoes your assessment predict anything? Validity for non-specialists
An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.
ReadSee a verified pipeline for one of your roles
Post a role free and review anonymous, skill-ranked candidates. No card, no sales call to get started.