A 4/5 Rating and an 82 AI Score Are Different Evidence
Compare what structured interview ratings and AI candidate scores can support, including validity, traceability, fairness, cutoffs and scale design.

A job-anchored structured interview rating is usually the stronger baseline; add an AI candidate score only when its role-specific value is demonstrated. Neither score predicts performance by itself. For federal structured-interview design, OPM recommends at least 3 proficiency levels, aims for 5–7 and labels at least 3; there is no equivalent universal meaning for an AI score of 82.
Choose the score source and decision use; the guide shows what evidence you need.
Score Evidence Decision Guide
Open each check to compare what the two numbers can support.
Meaning of the number
It should show how an answer matches predefined, job-related proficiency anchors.
It means only what the documented method defines: a trait estimate, rank, classification or prediction. It is not automatically a probability of job success.
Traceability and review
A reviewer can inspect the question, permitted probes, response notes, selected anchor and rater reasoning when those records are retained.
A label such as “communication: 82” does not reveal which evidence, features or transformations changed the score.
Job relevance
Questions and anchors can be tied to tasks and entry-level competencies identified through job analysis and subject-matter review.
Correlations or proxies do not by themselves establish that the model measures a necessary job characteristic.
Scale levels and apparent precision
There is no universal legal number of levels. OPM’s federal guide recommends at least 3 proficiency levels, aims for 5–7 and labels at least 3.
There is no universal 0–100 standard. Extra points or decimals may imply precision unsupported by validation evidence.
Reliability and validity
Test whether trained raters score the same evidence similarly. Greater structure is associated with greater rater reliability and agreement, but a sound rubric alone does not establish predictive validity.
Test consistency under relevant conditions, which may include transcription, device, language or retest variation. A vendor-wide claim does not establish validity for this role, population, outcome and use.
Fairness and accessibility
Common questions and anchors constrain discretion but do not eliminate biased content, unequal probing, inaccessible administration or rater error.
Using one model for everyone does not establish fairness. Inputs, labels, training data, missing data and prediction errors may affect groups differently.
Cutoffs and ranking
“3 means acceptable” is supportable only when the anchor represents acceptable job proficiency and raters apply it consistently.
“Advance everyone above 80” requires evidence for both the score interpretation and that operational threshold.
Monitoring and final choice
Prefer anchored ratings when interviews can elicit the required evidence and trained reviewers can evaluate it at the needed volume. Revisit questions as duties change, retrain raters and monitor scoring.
Add the score only when it contributes demonstrated value without replacing inspectable evidence. Applicant behavior, integrations, data and model versions can change outputs.
Sources: OPM; SIOP; NIST; 29 CFR Part 1607; EEOC.
The comparison draws on the federal OPM structured-interview guide, SIOP’s AI assessment recommendations, the NIST AI RMF Playbook, the U.S. Uniform Guidelines and EEOC selection-procedure guidance. Legal duties vary by tool, employer and jurisdiction.