Find the Reason Behind Conflicting Interview Scores
Trace conflicting interview ratings to evidence, criteria, opportunity or scale calibration before turning them into an average.

A candidate receives a two from one interviewer and a four from another. Averaging them to three produces a tidy number, but it does not explain the gap. Start with the evidence, the criterion and the scale. One of those may differ between reviewers.
Use the team’s existing job-related criteria and approved selection procedure. This check is about interpreting the evidence already collected, not introducing a new scoring system.
Symptom: the same answer has two scores
Have each reviewer identify the passage or observed action behind their rating. Then ask which line of the rubric it satisfies. Keep the initial ratings visible while you investigate; overwriting them immediately loses the distinction between a changed interpretation and a clerical correction.
If one reviewer heard “I resolved the request” and another recorded “the team resolved it,” the immediate issue is attribution. Consult the original record available under your process. Do not fill the gap with an AI summary that adds details or an interviewer’s confident recollection presented as a quotation.
Symptom: reviewers are scoring different things
One reviewer may reward speaking fluency while another rewards the quality of the decision. Both may have entered a number under “communication.” Rewrite the criterion into observable elements for future use and follow a consistent process for handling affected existing ratings.
OPM’s structured-interview description specifies common rating standards. Our inference is that a shared label without shared meaning is insufficient for comparison. OPM: Structured Interviews.
Symptom: one interviewer saw evidence the other did not
Separate “not observed” from “performed poorly.” If a follow-up question was asked in only one session, compare the opportunity candidates had to demonstrate that competency. Do not automatically turn a blank field into the lowest score.
The possible repair is a consistently administered clarification, not an improvised extra challenge for a candidate who happens to attract disagreement. Have the process owner decide whether more evidence is appropriate and how comparable treatment will be maintained.
Symptom: the scale differs between reviewers
Ask each reviewer what would earn the highest rating. If one reserves it for an impossible perfect answer and the other uses it for the defined standard, the issue is calibration. Try scoring a fictional response together, explaining the evidence at each step. Label the exercise as practice; do not pretend agreement proves the rubric predicts job performance.
Close the disagreement with a traceable decision
Record the original scores, the disputed evidence, the criterion applied and the reason for any change. If disagreement remains, keep it explicit for the authorized decision-maker. A short note such as “insufficient evidence on escalation judgment” is more informative than a compromise number with no explanation.
For the next interview cycle, repair the question, rubric or interviewer briefing that caused the mismatch. The useful outcome is a clearer assessment process, not simply unanimous ratings.
If the disagreement starts with the job itself, revisit the role requirements and assessment criteria.