calhire
All posts
IntegrityAssessmentIntegrityBiasAI in hiring

Why webcam proctoring is the wrong fix for assessment integrity

Webcam proctoring flags disabled candidates, people with darker skin and anyone without a private room. What it costs, and what to measure instead.

CThe CalHire TeamCalHire8 min read

Reviewed by CalHire Compliance, Compliance & Fairness

Webcam proctoring infers cheating from a candidate’s body and surroundings, so it systematically flags people who move differently, look different to the camera, or lack a private quiet room. It also introduces biometric processing and a surveillance experience, for signals that are weaker than session-level behavioural measures.

  • Proctoring flags proxies for disability, housing and skin tone — not for dishonesty.
  • Face and gaze analysis can constitute biometric or special-category processing, with extra legal conditions.
  • It shifts the assessment from "can you do the work" to "can you produce a compliant recording".
  • Session-level signals (latency, cadence, cross-candidate similarity) are stronger and far less invasive.
  • If you must use it, never auto-fail, always offer an alternative, and measure who gets flagged.

What proctoring is actually measuring

Remote proctoring watches a candidate through their camera and flags behaviour it treats as suspicious: eyes leaving the screen, a second face, a voice, a hand movement, a face that briefly fails to match, a window losing focus.

Read that list again and ask what it has in common. None of these are cheating. They are behaviours that correlate loosely with cheating in an idealised test-taking environment: one person, one quiet room, one still body, good lighting, uninterrupted.

The further a candidate is from that ideal, the more flags they generate. So the systematic question is not "does it catch cheats?" but "who does it flag?"

Who gets flagged

CandidateWhy they generate flags
A candidate with a motor or tic disorderInvoluntary movement reads as suspicious activity
A blind or low-vision candidate using a screen readerGaze does not track the screen as expected
A neurodivergent candidateLooking away to think, speaking aloud, stimming
A candidate sharing housingAnother person enters the frame
A parent or carerInterruptions are not optional
A candidate with a darker skin toneFace detection and matching accuracy has been shown to vary by skin tone and lighting
A candidate using a shared or public spaceBackground movement, noise, poor lighting
A candidate on an older laptop or weak connectionDropped frames read as absence or tampering

Every row is a proxy for disability, housing, income, caring responsibility, or skin tone. Not one is a proxy for dishonesty.

That is a textbook adverse-impact structure: a screen that removes groups at different rates for reasons unrelated to job ability. If you use proctoring and do not track flag rates by group, you are not measuring the thing most likely to matter — see the four-fifths rule.

There is a second-order cost too. A candidate who knows they will be flagged — because they always are — either declines to continue or sits the assessment under additional stress. Both show up as differential drop-off, which never appears in a selection-rate analysis at all.

The legal surface is genuinely awkward

Not automatically unlawful, and not simply fine either. Four things to work through before deploying it:

Biometric and special-category processing. Facial analysis or face matching may constitute biometric processing, which typically requires an additional lawful condition beyond an ordinary basis. Several jurisdictions have specific rules for biometric data and for AI analysis of video interviews.

Accommodation obligations. If the assessment format disadvantages a disabled candidate, you are into reasonable-accommodation territory. "The proctoring system flagged you" is not a defence, and the accommodation cannot be "we ignored the flags" if the candidate was still put through a distressing process.

Automated adverse decisions. If a proctoring signal ends someone's application, that is an automated decision about a person with significant effect. It carries oversight and explanation obligations under the EU AI Act and restrictions under GDPR.

Proportionality. The recurring question from data-protection regulators is whether a less intrusive means would achieve the same purpose. Given that session-level signals exist, that is a hard question to answer well.

What proctoring does to your assessment

The subtler cost, and the one that gets ignored: it changes what you are measuring.

An assessment is supposed to measure whether someone can do the job. Under proctoring it also measures whether they can produce a compliant recording — hold still, stay lit, keep the room clear, keep eyes forward, don't think out loud. For almost every role, none of that is job-relevant. You have added variance that is pure noise with respect to the thing you care about, and non-random noise at that, which is worse than random.

You have also told the candidate something about your company. Surveillance as a first interaction is a statement of the relationship you expect, and the strongest candidates have options.

What to do instead

1. Remove the incentive. Generate each assessment uniquely per session so there is no shared answer key. Most leakage is circulated answers, not real-time assistance.

2. Measure the session, not the person. Response latency patterns, typing cadence, paste behaviour, and cross-candidate answer similarity. These are proportionate, do not require a camera in someone's home, and are stronger evidence — cross-candidate similarity in particular is far more diagnostic than any gaze signal.

3. Verify identity once, rather than monitoring continuously. Confirming who is sitting the assessment is a defensible, bounded step. Watching them for ninety minutes is not the same thing, and does not follow from it.

4. Design tasks where assistance is not decisive. Critique over generation, trade-offs over recall, follow-up on their own submission. See candidates are using AI in your assessments.

5. Use a short human follow-up. Ten minutes asking someone to explain their own submission resolves almost every integrity question, in both directions, and is much cheaper than the false-accusation problem.

6. Never auto-fail. Any integrity signal should open a human review with a route for the candidate to respond.

If you have to use it anyway

Some organisations face client or regulatory requirements. Minimum guardrails:

  • No automated failures. Ever. Flags open reviews.
  • A genuine alternative route, offered proactively rather than buried in an FAQ.
  • Full disclosure of what is captured, analysed and retained, before the assessment starts.
  • Short retention with automated deletion, and no reuse of recordings for anything else.
  • No emotion or personality inference from video. The scientific basis is weak and the legal exposure is not.
  • Flag rates measured by group, reviewed quarterly, with action if they diverge.
  • A documented proportionality assessment you would be comfortable showing a regulator.

Where CalHire stands

We do not use webcam snapshot proctoring, and this is a stated design position rather than a roadmap gap.

The integrity layer instead combines identity assurance with liveness checks at the point of assessment, unique-per-session generation, and behavioural signals — response latency, typing cadence, LLM-likeness and originality — into a risk score that routes to a human reviewer, never an automatic verdict. No code path can auto-reject a candidate.

The AI interview is text-only; video is human-only and only after a consented identity reveal. And because personal information is stripped before any model sees a candidate, an integrity assessment cannot be influenced by who the candidate is.

Accommodations are documented rather than improvised, and every applicant receives a growth-framed report card regardless of outcome. See the features page for the integrity layer and for candidates for the candidate-side view.

The core error in proctoring is treating an assessment as a security perimeter to defend. It is a measurement instrument. Adding surveillance does not make a weak instrument stronger — it adds noise to it, and sends the bill to the candidates least able to absorb it.

Frequently asked questions

What is wrong with webcam proctoring?
It infers dishonesty from appearance and movement. Looking away, moving involuntarily, speaking aloud while thinking, having someone else in the room, or being poorly lit all generate flags — and none of them are evidence of cheating. The candidates most likely to be flagged are disabled candidates, neurodivergent candidates, people sharing housing, carers, and candidates whose faces the system handles less accurately.
Is webcam proctoring legal?
It depends heavily on jurisdiction and implementation. Facial analysis may constitute biometric or special-category processing requiring an additional lawful condition; several jurisdictions regulate AI analysis of video interviews specifically; and an assessment that disadvantages disabled candidates raises accommodation obligations. It is rarely simply prohibited and rarely simply fine. Take advice.
What should we use instead?
Signals about the session rather than the person: response latency patterns, typing cadence, whether content arrived pasted, and cross-candidate answer similarity. Combine them into a risk score that opens a human review. Then reduce the incentive to cheat by generating each assessment uniquely, so there is no shared answer key.
What if our clients or regulators require proctoring?
Then implement it with guardrails: never auto-fail on a proctoring signal, always offer a genuine alternative assessment route, disclose exactly what is recorded and retained, keep retention short, and measure flag rates by group. If the flag rate differs materially across groups, you have found a problem regardless of what the requirement says.
Share this post

Keep reading

Integrity8 min

Candidates are using AI in your assessments. Now what?

AI-text detectors are unreliable and disproportionately flag non-native speakers. What to do instead: assessment design, behavioural signals, and human review.

Integrity8 min

Fake candidates: proxy interviews, stolen identities and deepfakes

Remote hiring created a real identity-fraud problem: proxy interviewers, borrowed identities and deepfaked video. How the fraud works and where to verify.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.