HRaizon

Feature

Why a Confidence Score Is Not a Legal Explanation—or Proof the AI Is Right

By Priya Ellison ·

The short answer: confidence is a triage signal, not a legal shield

No. Adding a confidence score does not, by itself, make an AI recommendation explainable, correct, fair, or legally defensible. The materials reviewed for this article do not establish a universal legal rule requiring every legal, hiring, or other consequential AI recommendation to display a confidence score. They instead show that explanation and review obligations depend on the governing framework and context, while a numerical score may answer only a narrow technical question (see the legal-AI regulatory discussion).

A properly defined and validated confidence signal can still help an organization:

  • prioritize items for review;
  • identify borderline outputs;
  • select samples for quality control;
  • escalate unusual or uncertain matters;
  • request additional facts or research;
  • route an issue to a qualified professional; or
  • let the system abstain instead of forcing an answer.

Those are workflow functions. They are not proof that the underlying recommendation is accurate or legally sound.

The relevant duties may vary with the jurisdiction, decision, people affected, consequences of error, and degree to which the AI shapes the result. A document-prioritization tool presents different risks from a system that ranks applicants, recommends termination, predicts litigation outcomes, or determines access to a benefit.

A useful interface may display a score, explanation, citations, and proposed next action. But the interface does not establish that:

  • the score measures what users think it measures;
  • the recommendation is factually correct;
  • cited materials are authentic, current, or controlling;
  • the system was validated for the current population or matter;
  • the process treats relevant groups fairly;
  • a reviewer exercised independent judgment; or
  • the deployment satisfies applicable law.

The practical test is therefore not simply, “How confident is the AI?” Users also need to ask:

  1. What does the score measure?
  2. What evidence supports the recommendation?
  3. What assumptions, missing facts, or contrary sources could make it wrong?
  4. Who is responsible for reviewing, challenging, or overriding it?

Important: This article is informational only and is not legal, HR, or employment advice. It is not a current jurisdiction-by-jurisdiction survey. Laws and guidance vary and change frequently. Confirm requirements for the specific system, location, and decision with qualified counsel before acting, consistent with HRaizon’s advisory limitations.

Six concepts that should not be collapsed into one score

Confidence scoring becomes dangerous when one authoritative-looking number is asked to represent several different questions. At least six concepts should remain distinct.

1. Model confidence

Model confidence is a system-specific indication of preference, fit, relevance, or uncertainty. Depending on the system, it might represent the relative strength of one classification, similarity to retrieved material, predicted outcome probability, match strength, or an internally constructed score.

Unless the score has been defined and validated as a probability of correctness, users should not interpret it that way. A high value may mean only that the model strongly prefers one option over the alternatives it considered.

2. Calibration

Calibration describes the aggregate relationship between predicted scores and observed outcomes on evaluated data. If a score is intended to represent probability, calibration asks whether outputs assigned similar probabilities produce corresponding outcome rates over many cases.

Calibration is not individual verification. A well-calibrated system will still make errors, and calibration cannot tell a reviewer whether the recommendation currently on screen is one of them.

3. Verification

Verification means independently checking a specific claim, quotation, citation, fact, or conclusion against reliable evidence. In legal work, that can include opening an authority, confirming that it exists, reviewing the cited passage, checking subsequent treatment, and determining whether it applies to the relevant jurisdiction and facts.

Verification is item-specific. It asks whether this recommendation is supported—not merely whether the model performed acceptably across a past dataset.

4. Transparency

Transparency concerns information about the system and its operation. It may cover intended use, development methods, data sources, validation, limitations, update history, responsible owners, thresholds, and monitoring arrangements.

Transparency can help users understand when a system should or should not be trusted. It does not necessarily explain why a particular output occurred, and disclosure of architecture or training information does not automatically make an individual recommendation reviewable.

5. Explainability

Explainability is a human-usable account of the factors, criteria, evidence, or examples associated with an output. An explanation might identify the résumé criteria that affected a ranking, the contract clauses that triggered a risk flag, or the facts associated with a legal recommendation.

A post-hoc explanation can facilitate review without reproducing the model’s internal process. Feature importance, a generated rationale, or a counterfactual may be useful, but it should not be presented as a faithful transcript of machine “reasoning” unless that claim can be substantiated.

6. Legal justifiability

Legal justifiability is the ability to connect a recommendation to applicable facts, authorities, rules, precedents, contractual provisions, approved policies, or other legitimate criteria—and to expose that connection to challenge.

A legally useful explanation does more than say that an input had mathematical influence. It shows why the input matters under the governing standard. “Employment gap contributed to the score” is not, by itself, a legal or job-related justification. “The application does not contain evidence of a license identified as mandatory in the approved role criteria” is more reviewable, although the fact and criterion still need verification.

These concepts should also be separated from:

  • factual support;
  • source authenticity and quality;
  • evidentiary strength;
  • recommendation strength;
  • model uncertainty;
  • legal uncertainty; and
  • operational risk.

An expert-assigned legal-certainty scale illustrates the distinction. The Generative AI Legal Explainer uses a scale from 1, where the answer is effectively uncertain, to 5, where decades of case law support it. Longer answers describe legal context and facts that could change the result. That is a scale for how settled a general legal answer appears—not a model-generated probability that a particular outcome will occur (the project explains its scale here).

A multidimensional display is usually more informative than a composite percentage:

  • Model uncertainty: low
  • Factual support: incomplete
  • Source quality: mixed
  • Legal certainty: unsettled
  • Recommended action: qualified review required

The profile is less tidy than one number, but it tells users where the actual weaknesses are.

Why high confidence can coexist with a wrong recommendation

Many model scores are relative rather than evidentiary. In a typical classification system, a softmax function converts model outputs into values between zero and one that sum to one. A value such as 0.97 indicates strong preference among the alternatives evaluated; it does not inherently establish a 97% probability that the answer is factually or legally correct. The score comes from the model’s processing, not an independent investigation of reality (this analysis distinguishes softmax preference, calibration, and verification).

Even strong aggregate calibration cannot identify whether a particular output is wrong. Evaluation across thousands of past items may support a bounded performance claim under the test conditions. It cannot establish that a particular quotation is accurate, an applicant meets a criterion, or a cited case supports the proposition attributed to it.

One score may also obscure different forms of uncertainty:

  • Aleatoric uncertainty arises from ambiguity or irreducible noise in the data. Two qualified reviewers might reasonably classify an unclear document differently.
  • Epistemic uncertainty arises from lack of knowledge, such as unfamiliar facts, a novel legal issue, missing coverage, an incomplete source collection, or an input unlike those used in validation.

A system may appear confident because it does not recognize its own ignorance. Those situations require different responses, yet one percentage may not reveal which is present.

A bounded legal example shows the verification gap. In 2023, lawyers submitted a filing in a New York federal court containing nonexistent cases generated by a chatbot. The incident does not prove that every legal AI system behaves the same way. It does show why professional users cannot treat polished presentation or plausible citations as authentication (the incident is summarized in this confidence-and-verification analysis).

A common failure pattern looks like this:

  1. An AI system recommends that a motion is unlikely to succeed.
  2. It assigns the recommendation a high score.
  3. It produces a polished explanation using familiar legal language.
  4. It lists apparently relevant authorities.
  5. One authority is fabricated, or a real case does not support the stated proposition.
  6. The recommendation rests materially on that defective authority.

The score, prose, and citation formatting reinforce one another psychologically. Precise percentages can appear scientific. Fluent language can appear expert. Citations can create the impression that research has already been completed.

Citations improve inspectability only when a user opens and evaluates the underlying authority. The reviewer should check:

  • Authenticity: Does the source exist, and is the quoted text accurate?
  • Relevance: Does it address the proposition at issue?
  • Jurisdiction: Is it binding, persuasive, or inapplicable?
  • Currency: Has it been amended, reversed, superseded, or distinguished?
  • Completeness: Are contrary authorities or limiting passages missing?
  • Application: Do the rule and facts actually support the recommendation?

External methods can supplement model confidence. Retrieval checks can compare claims with source documents. A separate fact-checking layer can assess whether a source supports a generated statement. Independent reviewers can verify decisive propositions. Conformal methods can produce prediction sets with statistical coverage under applicable assumptions.

None of those controls guarantees that an individual legal answer is correct. Retrieval depends on the quality and completeness of the sources. A checking layer may confirm consistency with an incomplete collection. Statistical coverage does not authenticate a quotation or resolve a disputed interpretation. These methods answer narrower questions than “Is this recommendation right?”

What an explainable AI recommendation should actually show

An explainable recommendation should let a competent person inspect the path from governing criteria and evidence to proposed action. It should also identify uncertainty and state what happens next.

The following is an illustrative governance template, not a universal legal form:

Field What it should contain
Recommendation The proposed conclusion or action, without overstating certainty
Intended decision The workflow or decision the output is meant to support
Score definition What the score measures, how it may be used, and prohibited interpretations
Supporting facts Material facts relied on and their sources
Governing criteria Applicable law, policy, rubric, contract term, playbook rule, or approved standard
Criterion-level analysis Evidence associated with each decisive criterion
Sources Documents, statutes, cases, policies, records, or other authorities
Contrary evidence Conflicting facts, authorities, findings, or plausible alternatives
Assumptions Conditions treated as true but not independently established
Limitations Missing information, weak coverage, out-of-scope issues, or model constraints
Uncertainty profile Separate assessments of model uncertainty, legal certainty, factual support, and source quality
Change conditions Facts, authorities, or inputs that could alter the recommendation
Required reviewer action Verify, request information, escalate, approve, override, or abstain
Responsibility and challenge route Who owns the decision and how review may be requested

Where practical and permitted, material claims should lead directly to the supporting authority or evidence. A legal proposition should point to the relevant statute, regulation, case, contract provision, policy, or record—not merely a search page or generated bibliography. A hiring assessment should identify the approved job criterion and the candidate evidence associated with it, not only a composite “fit” score.

Decisive criteria should remain separate. If a contract recommendation depends on governing law, liability-cap language, and a playbook threshold, users should be able to see how each criterion was treated. If a candidate recommendation depends on a required credential, relevant experience, and structured responses, those elements should also remain visible.

The record should disclose what could change the result, including:

  • a missing document or incomplete application;
  • contradictory evidence;
  • a different jurisdiction or governing-law clause;
  • authority from the wrong date;
  • an unsettled legal issue;
  • a weak or secondary source;
  • updated legislation, policy, or case law;
  • an input outside the system’s validated scope; or
  • an unconfirmed factual assumption.

Natural-language rationales may make review easier, but they must be described carefully. A rationale can summarize factors associated with an output without revealing the model’s actual causal process. Feature attribution and post-hoc explanations can create an appearance of understanding while leaving substantial gaps in operator knowledge and governance (Jones Walker discusses these limitations).

Whenever possible, replace an unexplained percentage with a practical profile:

Recommendation: Escalate the clause for specialist review. Model uncertainty: Low. Factual support: Complete for the supplied contract. Source quality: Current internal playbook; no jurisdiction-specific authority retrieved. Legal certainty: Not assessed. Next action: Qualified review required before advice is given.

Affected people may also need recipient-facing information. Depending on the context and applicable requirements, that can include:

  • whether AI materially influenced the process;
  • the principal criteria applied;
  • who remains responsible for the decision;
  • how to request human review;
  • how to correct inaccurate information;
  • how to challenge the result; and
  • what additional evidence may be submitted.

An explanation is useful when it enables action, verification, and challenge—not merely when it makes the output sound reasonable.

Turn uncertainty into a review workflow, not an automatic decision

There is no universal numerical threshold at which an AI recommendation becomes safe to accept, reject, or automate. Thresholds should be tied to a defined score, validation evidence, error costs, the decision context, and applicable duties.

The following matrix is illustrative and deliberately non-numerical:

Situation Appropriate response
Validated high confidence, strong current sources, familiar matter Prioritize the item; retain risk-based sampling and any required professional review
Borderline recommendation Review decisive evidence, compare it with governing criteria, and document the reviewer’s conclusion
Low confidence or insufficient information Obtain missing inputs, conduct further research, escalate, or abstain
High model confidence but weak, old, incomplete, or conflicting sources Let source risk override the score; require verification and qualified review
Low model confidence but strong direct authority Investigate why the model and evidence diverge; do not automatically reject the source-supported conclusion
Novel issue or input outside validated scope Abstain or route the matter to a specialist
Material disagreement between systems or reviewers Escalate and preserve the competing analyses

A high score may justify faster routing, but not necessarily less scrutiny. Serious consequences may still require professional review.

Borderline results need targeted review rather than ceremonial approval. The reviewer should examine the evidence that controls the outcome, compare it with the approved rule or rubric, record whether the recommendation was accepted or changed, and explain material overrides.

Low confidence should trigger an action—not merely a warning icon. Appropriate responses include requesting missing facts, broadening research, obtaining another source, escalating to counsel or an experienced HR professional, or allowing the system to return “insufficient information.”

Source risk should override model confidence. An outdated policy, incomplete record, secondary legal summary, or conflicting authority cannot be repaired by a strong model score. Conversely, low model confidence should not defeat clear direct evidence. A reviewer should investigate the divergence.

Meaningful human oversight requires more than an “approve” button. A reviewer should:

  • have the competence to evaluate the subject;
  • receive the relevant evidence and criteria;
  • understand the score and system limitations;
  • have adequate time to investigate;
  • possess real authority to override or stop the process;
  • avoid incentives that make approval automatic; and
  • remain accountable for the resulting action.

Nominal approval does not necessarily neutralize the system’s influence. If a ranking determines which applicants are seen, which documents are reviewed, or which matters receive attention, it may shape the outcome before a person makes the formal decision.

Legal document review offers a practical example. Scores can prioritize documents, identify borderline items for escalation, and support sampling or quality control. Rationales and citations can help reviewers compare a coding suggestion with the text. Practitioner-oriented product guidance nevertheless pairs those features with attorney review, sampling, and audit planning rather than treating the score as a final coding decision (see this legal-review workflow discussion).

Validate the score before anyone relies on it

A confidence score should not enter a consequential workflow without a written definition. At minimum, the definition should identify:

  • the outcome or construct being measured;
  • the unit of analysis, such as a document, candidate, claim, or matter;
  • the calculation method;
  • the intended users;
  • the represented population and data;
  • the relevant jurisdiction and time period;
  • the decisions the score may support;
  • conditions under which it should not be used; and
  • prohibited interpretations, especially “probability of correctness” if that meaning has not been validated.

Ask what the number actually represents. Is it relative model preference, match strength, relevance, predicted probability, retrieval similarity, agreement with a rubric, or something else? If the vendor cannot provide a precise answer, users cannot responsibly set thresholds or explain the score.

Validation should address the intended workflow, not merely overall accuracy. Evaluate calibration on representative data and near the operational points where recommendations are approved, escalated, sampled, or withheld.

Useful documentation may include:

  • sample size and selection method;
  • outcome definitions;
  • false-positive and false-negative rates;
  • common error types;
  • missing-data behavior;
  • out-of-distribution behavior;
  • confidence intervals or other measurement uncertainty;
  • known limitations;
  • consequences of changing thresholds; and
  • differences between testing and live use.

Performance should be examined across relevant matter types, document classes, jurisdictions, time periods, and demographic groups—not only in aggregate. A system may perform well overall while being overconfident, underconfident, or less accurate for a subgroup or category.

Fairness analysis adds separate diagnostics:

  • Demographic parity compares positive-outcome rates across groups.
  • Equalized odds compares true-positive and true-negative performance.
  • Equality of opportunity focuses on true-positive rates.
  • Group calibration asks whether similar scores correspond to similar observed outcomes across groups.

These metrics are not interchangeable legal tests. They can conflict, and satisfying one does not establish fairness in every relevant sense. Vendor-authored educational guidance provides a useful overview of these distinctions but should not be mistaken for a legal standard (see the fairness-metrics overview).

A statistical disparity may justify investigation, but it does not by itself establish its cause or amount to a legal finding of discrimination. Investigation may need to examine data quality, criteria, labels, proxy variables, accessibility, workflow design, reviewer behavior, and the applicable legal framework.

Validation is not a one-time procurement event. Monitoring should address:

  • data drift: live inputs differ from validation data;
  • source drift: retrieved collections become incomplete or outdated;
  • legal drift: statutes, regulations, cases, or policies change;
  • workflow drift: users change how they interpret or act on the score; and
  • threshold drift: operational cutoffs change without renewed testing.

Reevaluate the system after material changes to the model, prompt, rubric, source collection, workflow, threshold, population, or governing criteria. Independent evaluation and documented calibration evidence are stronger than unsupported vendor assurances, but neither guarantees correctness in an individual case.

Legal and hiring implications: explainability is context-specific

This section is issue spotting based on secondary materials, not definitive advice about any jurisdiction. It does not establish that a listed law applies to a particular tool or decision. Definitions, scope, effective dates, exemptions, amendments, implementing rules, and enforcement status should be checked against current official authorities.

Data protection and automated decisions

An academic preprint identifies GDPR Articles 13, 14, and 15 in connection with information about relevant automated processing, and Article 22 in connection with certain solely automated decisions producing legal or similarly significant effects. It also discusses human intervention and challenges to qualifying decisions while characterizing a generalized “right to explanation” as implicit rather than express. The authors emphasize that suitable explanations vary with the audience and context (review the preprint’s regulatory discussion).

That secondary discussion does not establish that a confidence percentage is a legally sufficient explanation. A number may fail to communicate the relevant logic, decisive factors, limitations, role of human review, or route for challenge. The supplied extract also does not support a detailed statement of EU AI Act obligations, so no such jurisdiction-specific conclusion should be drawn from it.

Professional obligations in legal practice

For legal professionals, relevant concerns may include competence, confidentiality, supervision, accountability, source verification, and preservation of professional judgment. Secondary commentary connects these concerns with ABA Model Rules 1.1, 1.6, 5.1, and 5.3 and recommends source-linked outputs, verification checkpoints, decision logs, and skeptical human review (see the legal-AI explainability discussion).

That commentary is not an official ethics opinion or substitute for the current rules adopted in the relevant jurisdiction. A lawyer should check applicable professional-conduct rules, court requirements, client duties, and ethics guidance. A vendor’s explanation feature does not determine whether professional obligations have been satisfied.

Hiring and employment decisions

In hiring, a recommendation is more reviewable when it can be traced to explicit, documented, job-relevant criteria. An unexplained fit score—or opaque use of tone, pace, eye contact, facial expression, or similar behavioral signals—raises basic governance questions: what is being measured, why it is relevant, whether it is valid for the role, and whether it creates or amplifies disparities.

A criterion-level explanation can identify:

  • the approved criterion;
  • why the criterion is relevant to the role;
  • the candidate evidence associated with it;
  • missing or conflicting information;
  • how the recommendation affected the process; and
  • who made the final decision.

Structured rubrics can improve inspectability, but they are not automatically valid, unbiased, or compliant. The criteria, evidence, application, and surrounding workflow still require assessment.

A November 2025 law-firm advisory describes a changing state and local patchwork. It reports that New York City Local Law 144 addresses bias audits, publication of audit summaries, and advance notice for covered automated employment decision tools. It also summarizes Illinois notice-related controls and Colorado impact-assessment, transparency, documentation, and appeal provisions for covered systems. The advisory warns that human involvement may not remove obligations when a ranking or score materially influences a decision (review Akerman’s hiring-law overview).

Those summaries are not substitutes for statutes, regulations, agency guidance, or current enforcement materials. Coverage may turn on detailed definitions, including whether the tool is regulated, whether it makes or substantially assists a decision, where the parties are located, and which employment stage is involved.

Human involvement should therefore be examined functionally. Ask whether reviewers see candidates excluded by the system, whether rankings determine attention, whether overrides occur in practice, and whether reviewers have enough information and authority to disagree. Final sign-off does not necessarily erase the system’s earlier influence.

Fairness testing and legal analysis also serve different purposes. Statistical metrics can reveal outcome or error-rate disparities and help target an investigation. They do not, standing alone, establish why a disparity occurred or whether it is unlawful. Legal conclusions require applicable law, evidence, context, and qualified analysis.

An audit and procurement checklist for confidence-scored AI

Procurement should begin with a written use case, not a product demonstration. Buyers need enough information to determine what the score means, whether it works in the intended context, and how failures will be detected and handled.

The following checklist is an illustrative governance framework. Particular controls may be required, optional, inappropriate, or insufficient depending on the deployment and applicable law.

Score definition and validation

Ask the vendor to provide:

  • the exact score definition;
  • the target outcome or construct;
  • the unit of analysis;
  • the calculation method;
  • whether the number represents probability, preference, similarity, relevance, or another quantity;
  • the calibration target and method;
  • representative validation data;
  • sample sizes and evaluation periods;
  • performance near proposed thresholds;
  • false-positive and false-negative rates;
  • subgroup and category-level results;
  • known limitations;
  • missing-data and out-of-distribution behavior;
  • prohibited uses and interpretations; and
  • independent evaluation results, if available.

“Proprietary” is not a complete answer when withheld information prevents responsible interpretation or oversight. Legitimate intellectual-property protections may constrain disclosure, but deployers still need enough evidence to govern the intended use.

Source and evidence controls

Ask:

  • Which sources can the system retrieve?
  • How is provenance recorded?
  • Can material claims be traced to supporting passages?
  • How are citations authenticated?
  • How are jurisdiction and date handled?
  • What happens when sources conflict?
  • How does the system identify missing evidence?
  • Can it distinguish primary from secondary authority?
  • How and when are source collections updated?
  • Can it abstain when evidence is insufficient?

A citation validator should not merely confirm that matching text exists. Reviewers must still determine whether the source proves the proposition for which it is cited.

Versioning and reconstruction

Consider requiring version records for:

  • model;
  • prompt;
  • system instructions;
  • rubric or playbook;
  • source collection;
  • retrieval configuration;
  • score calculation;
  • operational threshold; and
  • workflow rules.

The purpose is operational reconstruction: identifying which configuration generated a recommendation and what evidence was available at the time. Reconstruction need not reproduce every internal neural operation, but it should allow an auditor to understand the inputs, criteria, sources, output, and human actions.

Audit records

Subject to legal, privacy, privilege, retention, and security requirements, records may include:

  • input data;
  • retrieved sources;
  • recommendation;
  • score and score definition;
  • uncertainty and source-quality flags;
  • criteria applied;
  • system and prompt versions;
  • threshold in effect;
  • reviewer identity;
  • reviewer action;
  • override rationale;
  • final outcome;
  • challenge or appeal;
  • incident details; and
  • remediation.

Logging is not risk-free. Records may contain personal data, confidential business information, privileged material, sensitive employment information, or security-relevant details. Access, retention, deletion, preservation, and disclosure rules should be designed accordingly.

Workflow controls

Depending on the use case, request:

  • configurable approval gates;
  • exception escalation;
  • risk-based sampling;
  • system abstention;
  • human overrides;
  • second review for borderline matters;
  • alternative processes where appropriate or required;
  • appeal or challenge support;
  • exportable logs; and
  • controls preventing unauthorized threshold changes.

Vendor-authored recruiting guidance commonly recommends configurable human review, explanations, privacy controls, bias testing, security evidence, and retrievable audit logs. Those are useful procurement questions, but claims about a particular vendor’s product still need independent verification (see this recruiting procurement framework).

Ownership and accountability

Identify:

  • who owns the criteria;
  • who approves the score definition;
  • who selects thresholds;
  • who can change them;
  • who monitors results;
  • who investigates incidents;
  • who responds to affected people; and
  • who remains accountable when a vendor supplies the model or recommendation.

A third-party contract does not answer these governance questions. Responsibilities should be allocated explicitly, while recognizing that contractual allocation may not determine duties imposed by applicable law.

Privacy, confidentiality, and security

Review:

  • data minimization;
  • lawful collection and use;
  • candidate or client notices;
  • privilege and confidentiality handling;
  • encryption and access controls;
  • data lineage;
  • retention and deletion;
  • cross-border transfers;
  • use of customer data for model improvement;
  • incident response;
  • subprocessors; and
  • audit rights.

The amount of logging should be proportionate. “Log everything forever” may create unnecessary privacy, discovery, confidentiality, and security exposure.

User comprehension and automation bias

Test the interface with intended reviewers. Determine whether they can accurately answer:

  • What does the score mean?
  • What does it not mean?
  • Which evidence controls the recommendation?
  • What conditions require escalation?
  • Can they identify an intentionally planted error?
  • Do they override the system when appropriate?
  • Does a precise percentage produce more deference than a qualitative label?
  • Does the explanation clarify uncertainty or merely sound persuasive?

Measure actual reviewer behavior, not only training completion. Oversight is weak if reviewers have override authority on paper but lack time to investigate, rarely detect errors, or are penalized for disagreement.

Deployment and recurring governance

A prudent deployment rule is:

Do not automate a consequential action from a confidence score unless the specific use, score definition, threshold, validation evidence, oversight design, affected population, and legal basis have been assessed.

After deployment:

  • monitor errors, complaints, and overrides;
  • review incidents and near misses;
  • reassess calibration and subgroup performance;
  • sample high-confidence outputs;
  • inspect disagreements between reviewers and the system;
  • retrain reviewers;
  • update sources and governing criteria;
  • test after material changes; and
  • suspend or narrow the use when evidence no longer supports it.

Frequently asked questions

Is an AI confidence score the probability that a recommendation is correct?

Not necessarily. It may represent relative model preference, similarity, relevance, match strength, or another system-specific quantity. It should not be described as a probability of correctness unless that interpretation has been defined, calibrated, and validated for the intended use.

Even a calibrated probability describes aggregate outcomes under particular conditions. It does not independently verify a specific recommendation, fact, quotation, or citation.

Does the law require AI recommendations to include confidence scores?

The reviewed materials do not establish a universal requirement that every AI recommendation display a confidence score. Vendor guidance generally presents confidence indicators as design and governance tools for communicating uncertainty—not as universal legal mandates (see this confidence-indicator design discussion).

Particular regimes may instead require notices, information, documentation, audits, human review, impact assessments, or opportunities to challenge certain decisions. Whether any requirement applies depends on the jurisdiction, tool, decision, affected person, and degree of automation or influence. A numerical score is neither automatically required nor necessarily sufficient.

Can a plausible AI explanation be misleading even when it cites contributing factors?

Yes. A generated rationale may identify inputs associated with an output without faithfully reproducing the model’s internal processing. Feature attribution may indicate influence without proving causation, factual correctness, or legal relevance. Fluent wording can make an incomplete or post-hoc explanation sound more authoritative than it is.

Treat the explanation as a review aid. Check the evidence, criteria, source links, assumptions, omitted factors, and contrary information before relying on it.

What should a human reviewer verify before accepting an AI-generated legal recommendation?

The reviewer should verify:

  • material facts;
  • decisive quotations and citations;
  • authenticity and current status of the authorities;
  • jurisdiction and precedential weight;
  • whether each cited passage supports the stated proposition;
  • contrary authority;
  • factual assumptions and missing information;
  • the applicable rule, policy, or professional standard;
  • the score’s actual definition;
  • whether the matter falls within the validated use; and
  • whether the conclusion follows from the verified evidence.

The reviewer should also record material corrections, overrides, unresolved uncertainty, and the basis for the final judgment. Secondary legal-AI commentary similarly recommends source-linked outputs, verification checkpoints, and decision logs, but applicable professional rules and ethics guidance should be checked directly.

Does keeping a human in the loop make an AI-assisted hiring decision compliant?

No—not automatically. Human involvement is relevant, but an AI ranking or score may materially influence who is reviewed, interviewed, advanced, or rejected even when a person formally makes the final decision. Secondary employment-law analysis specifically warns that human sign-off may not remove obligations where the tool substantially influences the process.

Meaningful oversight requires a competent reviewer with relevant evidence, adequate time, knowledge of system limitations, authority to override, and accountability. Employers must separately assess any applicable notice, audit, accommodation, documentation, impact-assessment, and appeal obligations for the particular jurisdiction and tool.

The practical rule is simple: never ask only how confident the AI is. Ask what the score measures, whether it was validated for this use, which evidence supports the recommendation, what uncertainty remains, and who can challenge or override it. Confidence can help direct attention. Verification, qualified judgment, contestability, monitoring, and documented accountability support stronger governance and reviewability—but no score or checklist guarantees that a particular decision is correct, fair, defensible, or compliant.