calhire
All posts
ComplianceComplianceAI in hiringAssessment

What counts as an automated employment decision tool?

The AEDT definition turns on whether a tool substantially assists or replaces discretionary decision-making. Where the line falls, and the cases teams get wrong.

CCalHire ComplianceCompliance & Fairness7 min read

An automated employment decision tool, or AEDT, is a computational process derived from machine learning, statistical modelling, data analytics or artificial intelligence that issues a simplified output — a score, classification or ranking — used to substantially assist or replace discretionary decision-making about hiring or promotion.

  • Three elements: a computational process, a simplified output, and substantial assistance to a decision.
  • "Substantially assist" is the contested element — a score that is one of several equally-weighted inputs may fall outside it, but a ranking reviewers follow does not.
  • Calling something a "recommendation" does not exempt it if reviewers in practice defer to it.
  • A rules-based keyword filter can still be in scope depending on how it is built and used.
  • Whether a tool is an AEDT is a question about your workflow, not just about the vendor’s technology.

This is general information, not legal advice. Definitions and interpretive guidance vary by jurisdiction and change over time. Take advice on your specific tools and workflow.

The three-part test

The term comes from New York City's Local Law 144, and it has spread into general usage. Stripped to its structure, an AEDT has three elements:

  1. A computational process — derived from machine learning, statistical modelling, data analytics or AI.
  2. That issues a simplified output — a score, a classification, a ranking, a recommendation.
  3. Used to substantially assist or replace discretionary decision-making about employment.

All three must be present. Element three is where every interesting argument happens.

Element 1: "computational process" is broader than "AI"

People read this element as "does it use machine learning?" It does not say that. It says a computational process derived from machine learning, statistical modelling, data analytics or artificial intelligence.

Statistical modelling and data analytics are doing a lot of work in that sentence. A scoring formula someone derived from historical hiring data is a computational process derived from data analytics, even if it now runs as twenty lines of arithmetic with no model in sight.

The practical implication: you cannot exit the definition by simplifying the technology. A team that replaces a model with a hand-tuned point system built from the model's behaviour has changed the implementation, not the character.

Element 2: the "simplified output"

The output has to be a reduction — a number, a band, a rank, a yes/no, a shortlist. This is what distinguishes an AEDT from a tool that merely surfaces information.

Things that are typically simplified outputs:

  • A match or fit score
  • A percentile or ranking within a candidate pool
  • A pass/fail or above/below-threshold classification
  • An ordered shortlist
  • A risk or integrity flag that gates progression

Things that typically are not, on their own:

  • A transcript
  • A keyword highlight that leaves the document intact for a human to read
  • A calendar or workflow automation with no evaluative output
  • Search results a recruiter composes their own query for

The distinction is whether the tool has collapsed a judgement into a value, or presented information for a human to judge.

Element 3: "substantially assist or replace" — the real question

This is a question about your workflow, not the vendor's technology. Two companies using the same product can land on different sides of it.

The fact pattern most likely to fall outside the definition is an output that is one of several inputs, weighted no more heavily than the others, and not relied on as the primary criterion. The fact patterns clearly inside it:

  • The tool ranks and reviewers work top-down until the shortlist is full.
  • A threshold removes candidates before a human sees them.
  • A score is the only quantitative input and everything else is impressionistic.
  • The interface makes following the recommendation one click and deviating from it a form.

Labels do not help you here. "It's only a recommendation, a human always decides" is a claim about process that has to survive contact with the data. If 98% of decisions match the tool's ranking, and reviewers spend eleven seconds per candidate, the human is ratifying rather than deciding. That is the substance regulators look at, and it is also the substance of the human-oversight requirement under the EU AI Act.

The uncomfortable corollary: you should measure your own override rate. If you cannot say how often your reviewers disagree with the tool, you cannot support the claim that they are deciding.

Boundary cases teams get wrong

CaseCommon assumptionBetter analysis
Simple keyword filter"Too basic to be AI"If it eliminates candidates pre-human-review, it is exercising decision authority
Vendor "recommendation" label"Recommendations are exempt"Function over label; check actual deference
Sourcing / outreach tools"Not a hiring decision"Targeting who sees a job ad is squarely covered in some regimes, including the EU AI Act's Annex III
Integrity or fraud flags"That's security, not selection"If a flag gates progression, it affects the employment decision
Assessment scoring"It's just a test score"A score used to rank or threshold is the paradigm case
Chat-based screening"It's a conversation"If it emits a rating, it emits a simplified output

What to do with an answer of "yes, probably"

Being an AEDT is not a problem. Being an unexamined one is. If a tool is in scope:

  1. Write down the workflow, including exactly where the output enters the decision and who can override it.
  2. Measure the override rate, so your description of human involvement is evidence rather than assertion.
  3. Get the audit and publish it where required — see the Local Law 144 checklist.
  4. Notify candidates in advance, and build the alternative-process route before you need it.
  5. Keep configuration history, so you can show what the tool was doing on any given date.
  6. Check adverse impact yourself, on your own data — a vendor audit of a default configuration is not a statement about yours. See the four-fifths rule.

Where CalHire sits

CalHire is unambiguous about this: the platform produces a composite score and a ranking used in hiring, so we treat it as an automated employment decision tool and design accordingly rather than arguing about definitions.

That means: independent third-party bias audits published at /bias-audit; candidate notice before any automated tool is used, with a human-only alternative available on request; no code path that can auto-reject anyone, with below-threshold candidates flagged for human review; and every person-affecting decision logged, explainable, and tied to the human who made it.

It also means the identity of the candidate is not on the evaluation surface at all — personal information is stripped at an architectural boundary before scoring, the text interview, or ranking — so the class of bias that comes from a model seeing a name, photo, school or employer cannot arise. The design is described on the features page and the controls on the enterprise page.

If you are unsure whether a tool in your stack is an AEDT, the useful question is not "how advanced is it?" It is: if we switched it off tomorrow, would our shortlists change? If yes, it is assisting the decision.

Frequently asked questions

Is a keyword filter in our ATS an AEDT?
It can be. The definition turns on a computational process producing a simplified output used to substantially assist or replace discretionary decision-making — not on whether the technology is fashionable. A hard keyword filter that removes applications before any human sees them is exercising decision-making authority regardless of how simple the code is. Assess the workflow, not the sophistication.
Does labelling the output a "recommendation" avoid the definition?
No, and this is the most common misconception. What matters is the function the output serves in practice. If reviewers overwhelmingly follow the ranking, or if the interface makes deviating from it effortful, the tool is substantially assisting the decision whatever the label says. Regulators look at behaviour, not nomenclature.
What if the score is only one input among several?
That is the intended shape of the boundary — an output weighted no more heavily than other inputs, and not relied on as the primary criterion, is the fact pattern most likely to fall outside "substantially assist". But you have to be able to demonstrate it: show the other inputs, their weight, and cases where the human decision diverged from the tool.
Does an AI interview count?
If it produces a score, rating or ranking that feeds a hiring decision, treat it as in scope. If it merely transcribes and a human reads the transcript with no generated score, the analysis is different. The presence or absence of a simplified output used in the decision is the hinge.
Share this post

Keep reading

Compliance7 min

A practical guide to Emiratization-compliant hiring

UAE mainland firms must reach 10% Emirati hires in skilled roles. Here is how to hit the target with verified talent, blind assessment and audit-ready reporting.

Integrity8 min

Candidates are using AI in your assessments. Now what?

AI-text detectors are unreliable and disproportionately flag non-native speakers. What to do instead: assessment design, behavioural signals, and human review.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.