Skip to content
HRaizon Subscribe

A 4/5 Rating and an 82 AI Score Are Different Evidence

Compare what structured interview ratings and AI candidate scores can support, including validity, traceability, fairness, cutoffs and scale design.

Share X in f
Priya Ellison

A job-anchored structured interview rating is usually the stronger baseline; add an AI candidate score only when its role-specific value is demonstrated. Neither score predicts performance by itself. For federal structured-interview design, OPM recommends at least 3 proficiency levels, aims for 5–7 and labels at least 3; there is no equivalent universal meaning for an AI score of 82.

Choose the score source and decision use; the guide shows what evidence you need.

Score Evidence Decision Guide

Baseline recommendationPreserve the question, response evidence, proficiency anchor and rater reasoning. Use the rating to support review only when the scored competency is linked to the job.

Open each check to compare what the two numbers can support.

Meaning of the number
Structured rating

It should show how an answer matches predefined, job-related proficiency anchors.

AI score

It means only what the documented method defines: a trait estimate, rank, classification or prediction. It is not automatically a probability of job success.

Require: the construct, range, intended interpretation and decision use. For AI, also identify the inputs and predicted outcome.
Traceability and review
Structured rating

A reviewer can inspect the question, permitted probes, response notes, selected anchor and rater reasoning when those records are retained.

AI score

A label such as “communication: 82” does not reveal which evidence, features or transformations changed the score.

Require: evaluated input, output, scoring and validation documentation, plus the model or configuration version.
Job relevance
Structured rating

Questions and anchors can be tied to tasks and entry-level competencies identified through job analysis and subject-matter review.

AI score

Correlations or proxies do not by themselves establish that the model measures a necessary job characteristic.

Require: a current mapping from every scored dimension to job requirements. Proxies for concepts such as hireability need construct-validity evidence.
Scale levels and apparent precision
Structured rating

There is no universal legal number of levels. OPM’s federal guide recommends at least 3 proficiency levels, aims for 5–7 and labels at least 3.

AI score

There is no universal 0–100 standard. Extra points or decimals may imply precision unsupported by validation evidence.

Require: definitions for every consequential level or band. Use only distinctions that raters or evidence can support.
Reliability and validity
Structured rating

Test whether trained raters score the same evidence similarly. Greater structure is associated with greater rater reliability and agreement, but a sound rubric alone does not establish predictive validity.

AI score

Test consistency under relevant conditions, which may include transcription, device, language or retest variation. A vendor-wide claim does not establish validity for this role, population, outcome and use.

Require: a study design suited to the intended inference and appropriate content-, criterion- or construct-related evidence. A stable score can still measure the wrong thing.
Fairness and accessibility
Structured rating

Common questions and anchors constrain discretion but do not eliminate biased content, unequal probing, inaccessible administration or rater error.

AI score

Using one model for everyone does not establish fairness. Inputs, labels, training data, missing data and prediction errors may affect groups differently.

Require: examination of selection outcomes, administration conditions and errors for relevant groups. Under U.S. federal law, disproportionate exclusion and less discriminatory alternatives can matter.
Cutoffs and ranking
Structured rating

“3 means acceptable” is supportable only when the anchor represents acceptable job proficiency and raters apply it consistently.

AI score

“Advance everyone above 80” requires evidence for both the score interpretation and that operational threshold.

Require: documentation of why the cutoff reflects required proficiency and review of candidates near it. Under the U.S. Uniform Guidelines, evidence sufficient for pass/fail use may be insufficient for ranking when ranking creates greater adverse impact.
Monitoring and final choice
Structured rating

Prefer anchored ratings when interviews can elicit the required evidence and trained reviewers can evaluate it at the needed volume. Revisit questions as duties change, retrain raters and monitor scoring.

AI score

Add the score only when it contributes demonstrated value without replacing inspectable evidence. Applicant behavior, integrations, data and model versions can change outputs.

Require: version control, monitoring intervals and response triggers. Combining two weak measures does not create one strong assessment.

Sources: OPM; SIOP; NIST; 29 CFR Part 1607; EEOC.

The comparison draws on the federal OPM structured-interview guide, SIOP’s AI assessment recommendations, the NIST AI RMF Playbook, the U.S. Uniform Guidelines and EEOC selection-procedure guidance. Legal duties vary by tool, employer and jurisdiction.