calhire
All posts
Hiring PlaybookAssessmentHiring processSkills

Structured interviews: the highest-return change most teams never make

Same questions, same order, same rubric, scored independently. Structured interviews are unglamorous, uncomfortable to adopt, and the best value in hiring.

CThe CalHire TeamCalHire8 min read

A structured interview asks every candidate the same job-related questions in the same order, scores each answer against a rubric written in advance, and has interviewers record scores independently before discussing. The structure is what makes candidates comparable — without it you are collecting impressions, not evidence.

  • Structure means four things: same questions, defined probes, a rubric written first, and independent scoring.
  • Score before you discuss. The first opinion spoken in a debrief anchors the room.
  • Unstructured interviews feel more informative than they are — rapport is not signal.
  • Two interviewers scoring independently beats four interviewers converging in a meeting.
  • The last stage in your process is your real selection criterion. If it has no rubric, nothing upstream matters.

What "structured" actually means

Not "we have a list of topics". Four specific things, all of them at once:

  1. The same core questions, in the same order, for every candidate for that role.
  2. Defined probes — the follow-ups you are allowed to use, decided in advance, so one candidate is not helped along while another is left to flounder.
  3. A rubric written before you meet anyone, with scoring bands and anchor examples, including what a mediocre answer looks like.
  4. Independent scoring: each interviewer records their score before any discussion.

Remove any one and you lose most of the benefit. The most commonly skipped is the fourth, and it is the cheapest to fix.

Why unstructured interviews feel so much better than they are

An unstructured interview is enormously satisfying. You follow your curiosity, the conversation flows, and you come away with a strong sense of the person. That feeling is the problem: the confidence it produces is not tracking accuracy.

Three mechanisms do most of the damage.

You ask different people different questions. So you never compared them. You compared your impression of candidate A answering question X to your impression of candidate B answering question Y.

Rapport gets scored as competence. Conversational ease is a real skill for some jobs and irrelevant for most. It is also unevenly distributed by culture, first language, neurotype, and how many interviews someone has sat this month.

The first few minutes set the rest. An early impression shapes which follow-ups you choose, how generously you interpret an ambiguous answer, and what you remember afterwards. You then experience the rest of the interview as confirmation.

None of this makes interviewers bad at their jobs. It makes unstructured interviews a poor instrument, in the same way that a badly built assessment can be reliable and still measure nothing — see does your assessment predict anything.

How to build one

Step 1 — Derive questions from the actual work. Take three to five skills from the job description and write one or two questions each. Two formats carry most of the weight:

  • Behavioural — "Tell me about a time you shipped something that broke in production. What did you do in the first hour?" Past behaviour, specific instance, not a hypothetical.
  • Situational — "You're on call, a payment provider is timing out on 30% of calls, and finance needs today's reconciliation. Walk me through the first thirty minutes." Judgement under realistic constraints.

Avoid brainteasers and questions about the candidate rather than the work. "Where do you see yourself in five years" measures how well someone has rehearsed answering it.

Step 2 — Write the rubric before you see a candidate. For each question, three or four bands with a concrete anchor:

Q: Production incident, first hour. 1 — Weak. Jumps to a fix with no diagnosis; no mention of communication or blast radius. 2 — Adequate. Diagnoses before fixing; mentions rollback; communication is an afterthought. 3 — Strong. Stabilises first, defines blast radius, communicates to a named stakeholder, then diagnoses. Distinguishes mitigation from root cause. 4 — Exceptional. All of the above, plus what they changed afterwards so the class of failure could not recur silently.

Writing band 2 is the hard part and the valuable part. Bands 1 and 4 are easy; nearly every real candidate lands between them.

Step 3 — Define the probes. For each question, list two or three permitted follow-ups. This stops the accidental generosity that makes candidates incomparable — helping the candidate you already like, and letting the one you do not run out of road.

Step 4 — Score independently, then discuss. Everyone submits scores before the debrief. This single change removes the anchoring effect where the most senior or most confident voice speaking first sets the room's baseline. Then discuss the disagreements, which is where the useful information is. Agreement needs no meeting.

Step 5 — Look at your final stage honestly. Whatever your last gate is, that is your real selection criterion. If the process is four structured rounds followed by an unscripted founder chat with veto power, then your selection method is an unscripted founder chat. Either give it a rubric or remove its veto.

What structure costs, honestly

It is worth naming the trade-offs, because pretending there are none is why adoption stalls:

  • Upfront work. A rubric per role, written once, revised occasionally. Real, and it amortises.
  • It feels rigid at first. Interviewers who pride themselves on reading people experience the constraint as a loss of skill. Some of that loss is the point.
  • Repetition. Asking the same questions twenty times is boring. Boring and accurate beats interesting and arbitrary.
  • It exposes disagreement. Independent scoring reveals that your team does not actually agree on what "good" means. Uncomfortable, and much better known than unknown.

What you get

Comparability, first — the ability to say candidate A scored higher than candidate B on the same evidence. Then a defensible record: what was asked, how it was scored, who decided. Then better feedback, because a rubric score converts directly into something specific to tell a candidate, which is the difference between useful feedback and a form letter. And finally, reduced exposure to the drift where irrelevant characteristics leak into a decision nobody documented.

How CalHire structures it by default

The text interview on CalHire is generated per session from the skills set for the role, scored against consistent criteria, and it runs under the integrity layer. Two properties matter here:

  • It is uniform by construction. Every candidate for a role is assessed on the same skill dimensions, so the comparability does not depend on an interviewer remembering to be consistent.
  • The interview never sees who the candidate is. Personal information is stripped before any model is involved, so the score cannot absorb identity signals. The interview is text-only; video is human-only.

The interview contributes to a composite score alongside the skills test and role-fit, with weights the employer sets per role. AI recommends and explains, and a person makes and is recorded for every decision — nothing is auto-rejected. See the features page for the full picture.

If you change one thing in your hiring process this quarter, make it independent scoring before the debrief. It costs nothing and it is the closest thing to free accuracy available.

Frequently asked questions

What makes an interview "structured"?
Four things together: every candidate gets the same core questions in the same order; the follow-up probes are defined in advance; each answer is scored against a rubric written before any candidate was seen; and interviewers record their scores independently before any group discussion. Drop any one of the four and comparability degrades.
Does structure make interviews robotic or cold?
It does not have to. Structure governs the questions and the scoring, not your tone. You can be warm, explain the format up front, and still ask everyone the same things. Candidates generally report structured processes as fairer, because the basis of the decision is visible.
How many interviewers do we need?
Fewer than most teams use. Two interviewers scoring independently against a rubric produce more reliable signal than four interviewers who talk first and score afterwards, because independent judgements do not contaminate each other. Adding people to a badly structured process adds correlated noise, not accuracy.
Should we still have an informal "meet the team" chat?
Yes — but be explicit that it is not a selection stage, and do not let its output into the decision. If a chat can veto a candidate, it is a selection stage, and it needs a rubric like every other one. Most teams either give it structure or accidentally let it become the whole process.
Share this post

Keep reading

Hiring Playbook8 min

Does your assessment predict anything? Validity for non-specialists

An assessment that feels rigorous can predict nothing. What validity and reliability mean, how to check yours, and the four ways hiring tests quietly break.

Hiring Playbook7 min

Why candidates get ghosted, and how to actually stop

Ghosting is almost never malice. It is what happens when closing the loop is optional, unowned and manual. The four structural causes, and the fix for each.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.