Skip to content
HRaizon Subscribe

Do Not Buy the Demo: An Evidence Checklist for AI Hiring Vendors

A vendor-neutral checklist for validating hiring assessments: job relevance, bias testing, accessibility, privacy, notices and monitoring.

Share X in f
Priya Ellison

A polished demo is not validation. Before deploying an AI-assisted assessment, require evidence that the tool measures something relevant to the job, works in your intended setting and can be operated lawfully in every applicable jurisdiction.

Use this checklist for scored interviews, résumé screening, games, tests and other systems that recommend, rank or screen candidates. It is informational, not legal or employment advice; have qualified counsel confirm the rules applying to your employer, locations and use case.

Define the deployment—not just the product

Document the proposed workflow before reviewing vendor evidence:

  • jobs, grades and locations covered;
  • candidates or employees affected;
  • data collected, inferred and retained;
  • output produced: score, rank, recommendation or automatic rejection;
  • cutoff or other decision rule;
  • people who can see, override or act on the output;
  • accommodation, challenge and manual-review routes; and
  • ATS, interview platform and other integrations.

A tool evaluated for graduate sales hiring is not automatically suitable for warehouse supervision. Evidence about an optional recruiter aid also does not validate using the same score as a knockout rule. UK government guidance likewise tells buyers to define the purpose, intended outputs, integration, human oversight and applicant-access requirements during procurement (UK Department for Science, Innovation and Technology, March 2024).

The evidence checklist

What to request What acceptable evidence should answer Warning signs
1. System and decision map What inputs become which features, scores and recommendations? Where do thresholds enter? Can the system reject someone, or does it inform a documented human decision? “AI-powered” descriptions without the actual scoring and decision path
2. Intended-use statement Which jobs, populations, languages, countries and hiring stages were evaluated? What uses are expressly unsupported? A claim that the assessment works for every role or workforce
3. Job-analysis record Which competencies or job requirements does each score measure, and how were those requirements established for the role? Desirable-sounding traits with no documented connection to the work
4. Validation report Who conducted the study, when, on what sample and against what outcome? Does it report methods, uncertainty, limitations and results relevant to the proposed use? Accuracy percentages, client logos, testimonials or an unexplained “validated” label
5. Subgroup results Selection rates, score distributions and relevant error rates by group; sample sizes; treatment of missing demographic data; and analysis at the job and decision stage where the tool will be used One aggregate fairness number, no denominators or results pooled across materially different jobs
6. Accessibility evidence What was tested, against which accessibility standard and version? What barriers remain? Is there an accommodation route or alternative that still measures the relevant ability? “WCAG compliant” without test scope, results or remediation details
7. Data-protection file What data is collected and why? What supports the lawful basis where required? What are the retention, deletion, training-reuse, subprocessor, transfer and rights-request arrangements? Indefinite retention, undisclosed secondary use or “consent” presented as a universal answer
8. Candidate communications Does the draft notice explain the tool’s role, data used, criteria assessed, accommodation contact, challenge route and required timing? A generic privacy policy that does not describe the assessment
9. Human-oversight design What do reviewers see? How are they trained? When is review mandatory, what supports an override and what is logged? “Human in the loop” without authority, time or information to question the score
10. Change and monitoring controls Are there version records, update notices, revalidation triggers, disparity monitoring, complaint handling, escalation and suspension criteria? Can the employer export outcome data? Silent model changes or contractual limits on examining outcomes

Test job relevance separately from bias

In the United States, a selection procedure can violate federal anti-discrimination law if it disproportionately excludes a protected group and the employer cannot justify it under the applicable legal standard. Under Title VII, the EEOC explains that a procedure producing disparate impact must be job-related and consistent with business necessity; it should evaluate skills associated with successful performance of the particular job. The agency also points employers to the Uniform Guidelines’ methods of test validation (EEOC, Employment Tests and Selection Procedures).

Request the full validation report, not only a vendor summary. Record:

  • the job and applicant population studied;
  • how the target construct was defined;
  • the criterion used, such as structured performance ratings;
  • the validity results and reported uncertainty;
  • the assessment version and scoring configuration;
  • exclusions, missing data and sample limitations; and
  • what the evidence does—and does not—support.

Then evaluate your own configuration and population. A vendor study is an input to the employer’s decision, not a substitute for it.

Bias evidence answers a different question. It can identify disparities without establishing job relevance, while a favorable aggregate result can hide a problem in one job or hiring stage. Require results at the level where decisions occur. For the mechanisms that can create disparities, see how AI hiring bias enters without explicit protected-trait fields.

Make accessibility a deployment gate

For US employers covered by the ADA, the EEOC says a test that screens out people with disabilities must meet the applicable job-relatedness and business-necessity requirements. Tests must be administered so that results reflect the factor being measured rather than an impairment, and reasonable accommodation may be required (EEOC guidance).

Ask the vendor to demonstrate the accessible route and explain how alternatives preserve the assessment’s purpose. Extra time may not address a barrier caused by speech recognition, keyboard interaction or inaccessible instructions. The employer—not the vendor’s default workflow—should retain responsibility for accommodation decisions. A practical candidate-side process appears in the guide to AI video-interview accommodations.

Apply a jurisdiction gate before signing

Do not treat one vendor compliance packet as global approval. Create a row for every place where the employer will use the system or where affected candidates and employees are located. The examples below were checked on September 22, 2026; rules and effective dates can change.

United States—federal baseline. Federal anti-discrimination rules can apply to tests and selection procedures whether or not the product is marketed as AI. EEOC guidance addresses disparate impact, job relevance, disability-related testing and accommodations (EEOC). State and local rules may add duties.

New York City. If the configured system is an automated employment decision tool covered by Local Law 144, an employer or employment agency may not use it unless a bias audit was conducted no more than one year before use, information about the audit is publicly available and required notices have been provided. The city states that notice must be given 10 business days before use (NYC Department of Consumer and Worker Protection). Obtain the audit, check whether it covers the relevant system and use, and assign responsibility for publication and notice. The audit does not by itself establish job relevance, accessibility or compliance elsewhere.

United Kingdom. The ICO says a data protection impact assessment should be completed before using recruitment AI, ideally during procurement. It also tells organizations to identify a lawful basis, document controller and processor roles, seek evidence that the provider has mitigated bias, give clear privacy information and limit unnecessary collection and retention (ICO, 6 November 2024). Telling candidates about processing is not the same as establishing a valid basis for it, as explained in candidate consent versus a privacy notice.

Check current original sources for every deployment location.

Put the evidence into the contract and pilot

Attach the approved system description, intended uses, version and controls to the contract. Address access to validation and audit materials, testing of the configured system, outcome-data exports, complaint investigation, material-change notifications and suspension rights. Allocate responsibility for retention, deletion, security incidents, candidate requests and subprocessors.

Before launch, run a limited pilot against predefined acceptance criteria. Confirm that:

  • integrations preserve inputs, scores and decision records correctly;
  • recruiters understand what the output means and does not mean;
  • accommodation, challenge and manual-review routes work;
  • notices reach the right people at the required time; and
  • subgroup outcomes can be assessed at the relevant decision point.

Set written no-go conditions. These might include missing job-relevance evidence, unresolved accessibility barriers, subgroup results without usable sample information, inability to provide required notices, prohibited secondary data use or refusal to disclose material model changes.

Approval should be dated and conditional. Name an owner, review interval and reassessment triggers, such as a model update, new role, changed cutoff, new data source, material disparity or candidate complaint. Procurement establishes the evidence baseline; post-deployment monitoring tests whether the assessment continues to deserve use.