calhire
ComplianceComplianceAI in hiringAssessment

What counts as an automated employment decision tool?

The AEDT definition turns on whether a tool substantially assists or replaces discretionary decisions. Where the line falls, and the cases teams get wrong.

CCompliance & Fairness7 min read

An automated employment decision tool, or AEDT, is a computational process derived from machine learning, statistical modelling, data analytics or artificial intelligence that issues a simplified output, a score, classification or ranking, used to substantially assist or replace discretionary decision-making about hiring or promotion.

  • Three elements: a computational process, a simplified output, and substantial assistance to a decision.
  • "Substantially assist" is the contested element, a score that is one of several equally-weighted inputs may fall outside it, but a ranking reviewers follow does not.
  • Calling something a "recommendation" does not exempt it if reviewers in practice defer to it.
  • A rules-based keyword filter can still be in scope depending on how it is built and used.
  • Whether a tool is an AEDT is a question about your workflow, not just about the vendor’s technology.

This is general information, not legal advice. Definitions and interpretive guidance vary by jurisdiction and change over time. Take advice on your specific tools and workflow.

The three-part test

The term comes from New York City's Local Law 144, and it has spread into general usage. Stripped to its structure, an AEDT has three elements:

  1. A computational process. Derived from machine learning, statistical modelling, data analytics or AI.
  2. That issues a simplified output. A score, a classification, a ranking, a recommendation.
  3. Used to substantially assist or replace discretionary decision-making about employment.

All three must be present. Element three is where every interesting argument happens.

Element 1: "computational process" is broader than "AI"

People read this element as "does it use machine learning?" It does not say that. It says a computational process derived from machine learning, statistical modelling, data analytics or artificial intelligence.

Statistical modelling and data analytics are doing a lot of work in that sentence. A scoring formula someone derived from historical hiring data is a computational process derived from data analytics, even if it now runs as twenty lines of arithmetic with no model in sight.

The practical implication: you cannot exit the definition by simplifying the technology. A team that replaces a model with a hand-tuned point system built from the model's behaviour has changed the implementation, not the character.

Element 2: the "simplified output"

The output has to be a reduction, a number, a band, a rank, a yes/no, a shortlist. This is what distinguishes an AEDT from a tool that merely surfaces information.

Things that are typically simplified outputs:

  • A match or fit score
  • A percentile or ranking within a candidate pool
  • A pass/fail or above/below-threshold classification
  • An ordered shortlist
  • A risk or integrity flag that gates progression

Things that typically are not, on their own:

  • A transcript
  • A keyword highlight that leaves the document intact for a human to read
  • A calendar or workflow automation with no evaluative output
  • Search results a recruiter composes their own query for

The distinction is whether the tool has collapsed a judgement into a value, or presented information for a human to judge.

Element 3: "substantially assist or replace", the real question

This is a question about your workflow, not the vendor's technology. Two companies using the same product can land on different sides of it.

The fact pattern most likely to fall outside the definition is an output that is one of several inputs, weighted no more heavily than the others, and not relied on as the primary criterion. The fact patterns clearly inside it:

  • The tool ranks and reviewers work top-down until the shortlist is full.
  • A threshold removes candidates before a human sees them.
  • A score is the only quantitative input and everything else is impressionistic.
  • The interface makes following the recommendation one click and deviating from it a form.

Labels do not help you here. "It's only a recommendation, a human always decides" is a claim about process that has to survive contact with the data. If 98% of decisions match the tool's ranking, and reviewers spend eleven seconds per candidate, the human is ratifying rather than deciding. That is the substance regulators look at, and it is also the substance of the human-oversight requirement under the EU AI Act.

The uncomfortable corollary: you should measure your own override rate. If you cannot say how often your reviewers disagree with the tool, you cannot support the claim that they are deciding.

Boundary cases teams get wrong

CaseCommon assumptionBetter analysis
Simple keyword filter"Too basic to be AI"If it eliminates candidates pre-human-review, it is exercising decision authority
Vendor "recommendation" label"Recommendations are exempt"Function over label; check actual deference
Sourcing / outreach tools"Not a hiring decision"Targeting who sees a job ad is squarely covered in some regimes, including the EU AI Act's Annex III
Integrity or fraud flags"That's security, not selection"If a flag gates progression, it affects the employment decision
Assessment scoring"It's just a test score"A score used to rank or threshold is the paradigm case
Chat-based screening"It's a conversation"If it emits a rating, it emits a simplified output

What to do with an answer of "yes, probably"

Being an AEDT is not a problem. Being an unexamined one is. If a tool is in scope:

  1. Write down the workflow, including exactly where the output enters the decision and who can override it.
  2. Measure the override rate, so your description of human involvement is evidence rather than assertion.
  3. Get the audit and publish it where required. See the Local Law 144 checklist.
  4. Notify candidates in advance, and build the alternative-process route before you need it.
  5. Keep configuration history, so you can show what the tool was doing on any given date.
  6. Check adverse impact yourself, on your own data, a vendor audit of a default configuration is not a statement about yours. See the four-fifths rule.

The vendor question worth asking

Vendors have an incentive to argue their way out of this category, and some will. A useful test when you are buying: ask whether the vendor treats its own product as an AEDT, and what follows from that answer.

We treat CalHire as one without argument, because the platform produces a composite score and a ranking used in hiring, and definitional debates about that would be self-serving. A vendor willing to say the same thing has probably built for the obligations. One that opens with why the definition does not apply to them has told you where their engineering effort went.

The test that settles it

If you are unsure about a tool in your own stack, the useful question is not "how advanced is it?"

It is: if we switched it off tomorrow, would our shortlists change?

If the answer is yes, the tool is assisting the decision, whatever it is called in the contract and however simple the code turns out to be. Everything in this piece follows from that one answer.

Frequently asked questions

Is a keyword filter in our ATS an AEDT?
It can be. The definition turns on a computational process producing a simplified output used to substantially assist or replace discretionary decision-making, not on whether the technology is fashionable. A hard keyword filter that removes applications before any human sees them is exercising decision-making authority regardless of how simple the code is. Assess the workflow, not the sophistication.
Does labelling the output a "recommendation" avoid the definition?
No, and this is the most common misconception. What matters is the function the output serves in practice. If reviewers overwhelmingly follow the ranking, or if the interface makes deviating from it effortful, the tool is substantially assisting the decision whatever the label says. Regulators look at behaviour, not nomenclature.
What if the score is only one input among several?
That is the intended shape of the boundary, an output weighted no more heavily than other inputs, and not relied on as the primary criterion, is the fact pattern most likely to fall outside "substantially assist". But you have to be able to demonstrate it: show the other inputs, their weight, and cases where the human decision diverged from the tool.
Does an AI interview count?
If it produces a score, rating or ranking that feeds a hiring decision, treat it as in scope. If it merely transcribes and a human reads the transcript with no generated score, the analysis is different. The presence or absence of a simplified output used in the decision is the hinge.
Share this post

Keep reading

Compliance8 min

What an AI bias audit actually measures

A bias audit is an outcome analysis, not a code review. What impact ratios are, what an audit cannot tell you, and how to read a summary you have been handed.

Compliance9 min

The EU AI Act and hiring: what recruitment teams are responsible for

AI used to recruit or evaluate candidates is high-risk under the EU AI Act. What that means for employers, what providers owe you, and what to ask vendors.

Hiring decided by proven skills

Create a free verified profile, or see how anonymous-first hiring works for your team.