Eightfold’s Passing Audits Measure Different Products and Evidence
Compare Eightfold’s 2026 matching and AI Interviewer audit ratios, missing demographic labels, synthetic test limits and NYC employer compliance duties.

Eightfold’s two published 2026 bias audits received passing opinions, but they do not establish that its products are bias-free. The Matching Model audit uses historical applicant data. AI Interviewer’s race/ethnicity and intersectional results use synthetic interviews—not observed scoring outcomes for real applicants in those groups. Every reported impact ratio exceeds 0.80, but these are scoring-rate comparisons, not hiring rates.
BABL AI Inc. prepared both audits. Neither certifies that an employer’s hiring process complies with every applicable law. Those limits appear in the Matching Model report and AI Interviewer report.
Choose the product you use; compare its reported ratios and evidence limits.
Audit Evidence by Product
Lowest reported ratio: 0.880. Unknown demographic labels were excluded, and aggregate results do not establish outcomes for your employer.
| Comparison | Ratio | Evidence |
|---|---|---|
| Gender | 0.962–1.000 | Historical |
| Race/ethnicity | 0.938–1.000 | Historical |
| Intersectional | 0.880 lowest | Historical |
| Age | — | Not quantified |
Ratios compare rates of scoring at or above the overall median—not hiring rates. — means not quantified in this audit.
Data Limits and Interpretation
Matching data: January 2024–December 2025. Excluded entries: 74,997,062 with unknown gender and 85,587,944 with unknown race/ethnicity. Counts overlap and are not necessarily unique people.
The lowest intersectional ratio is for Asian female applicants. A ratio above 0.80 is not a finding that every deployment is free of discrimination.
Source: BABL AI’s published 2026 Eightfold Matching Model and AI Interviewer audit summaries, linked in the article. No employer-specific outcomes are supplied.
Both Audits Passed, But Their Ratios Are Not Hiring Rates
BABL issued passing opinions for disparate impact, governance and risk assessment for both products. The Matching Model report is dated March 3, 2026, with a March 26 signature. AI Interviewer’s report is dated June 29, with a July 8 signature. The figures below come from the reports’ findings tables.
| Test | Matching Model | AI Interviewer |
|---|---|---|
| Gender | Female reference: 1.000; male: 0.962 | Male reference: 1.000; female: 0.990 |
| Race/ethnicity | 0.938–1.000 | 0.963–1.000, synthetic interviews |
| Gender × race/ethnicity | Lowest: 0.880, Asian female applicants | Lowest: 0.939, synthetic interviews |
| Age | Not quantified | “Above 40” reference: 1.000; “Below 40”: 0.902, inferred labels |
Both reports define the scoring rate as the proportion of a demographic group scoring at or above the overall median. The impact ratio compares that rate with the highest scoring rate in the relevant table. A ratio of 0.938 does not mean 93.8% of applicants were hired, or that the model is 93.8% accurate.
The federal four-fifths rule is a screening rule of thumb—not a legal definition of discrimination. Federal guidance expressly says smaller differences can still indicate adverse impact when statistically and practically significant.
Separately, NYC’s law requires an audit but does not prescribe a particular action based on its results. Anti-discrimination obligations still apply, as the NYC FAQ explains. A passing opinion therefore does not settle whether a particular employer’s use is lawful.
Matching Model Uses Historical Data With Large Label Gaps
The Matching Model assesses skills and experience against a particular role and produces a zero-to-five-star score in half-star increments. Client-defined job requirements can affect the output. Its audit covers data from January 2024 through December 2025, with testing conducted by Eightfold in January 2026 and reviewed by BABL.
The gender table contains about 23.8 million entries with known labels; the race/ethnicity table contains about 15.5 million. However, the report excludes 74,997,062 entries with unknown gender and 85,587,944 with unknown race/ethnicity from those respective calculations. These exclusions overlap: do not add them together or assume the counts represent unique people. The counts appear on pages 13–14 of the report.
Large labeled samples do not establish that unlabeled applicants experienced the same scoring patterns. Nor do the aggregate tables establish equal outcomes for every employer or job. The published ratios describe the analyzed entries, not the excluded population or every local deployment.
Use the PDF tables rather than relying solely on the webpage narrative. Eightfold’s matching results page incorrectly describes non-Hispanic White male applicants as sharing the lowest 0.880 ratio. The PDF table gives that group 0.882; the lowest reported intersectional ratio, 0.880, belongs to Asian female applicants.
AI Interviewer’s Race Comparisons Use Synthetic Candidates
In the audited configuration, AI Interviewer uses an LLM to generate interview questions, transcribes spoken responses, and uses another LLM to score the transcript against a human-editable rubric. Job descriptions, questionnaires and scoring criteria are configurable.
Its historical gender data covers 1,250 male and 395 female applicants, with another 1,235 excluded for unknown gender. The gender and age data span January–December 2025.
For race/ethnicity and intersectional testing, insufficient historical data was available. Instead, an LLM acted as candidates across 12 personas, with demographic information represented through generated candidate names alone. The configuration and data details appear on pages 10–15 of the AI Interviewer report.
BABL explicitly highlights the synthetic testing in an Emphasis of Matter on pages 3–4 and 7: synthetic interviews may differ from real operating conditions, increasing uncertainty. Its passing opinion remains unchanged. But those results should not be presented as evidence that real applicants of every race experienced equivalent interview scoring. They also do not establish accessibility for candidates with disabilities.
The age result is supplemental. NYC’s required comparisons concern sex, race/ethnicity and their intersections, and DCWP guidance says inferred demographic data cannot be used for the required bias audit. The report’s additional inferred-age comparison is neither a substitute for those required comparisons nor a comparison using verified age labels. Keep its “Above 40” and “Below 40” categories distinct from any broader claim about age discrimination.
Employer Reliance Depends on Product, Data and Deployment
Match the Audit to the Product and Configuration
A Matching Model audit does not cover AI Interviewer, much less every feature in Eightfold’s recruiting suite. Before relying on either report, request the model/version, deployment dates, rubric and configuration records for the tool actually used. Our algorithm-change records guide explains what to retain.
This distinction matters because client-defined job requirements affect matching outputs, while interview questionnaires and scoring criteria are configurable. A vendor-level aggregate result does not establish the result for a particular employer’s settings.
Confirm Eligibility to Rely on Pooled Data
NYC guidance permits multi-employer historical audits, but an existing user must have supplied historical data from its own use to the auditor. First-time users can rely on other employers’ data. Test data is permitted when historical data is insufficient, with its source and rationale disclosed.
Do not assume that publication of a vendor audit alone establishes that your employer can rely on it. Check whether the relevant report and its data meet the requirements for your use.
Complete the Separate NYC Employer Duties
For covered NYC use, the employer must ensure an audit within the preceding year, a public results summary and publication of the tool’s distribution date—the date the employer began using it. Required notices to NYC-resident candidates and employees must precede use by 10 business days.
Neither report determines whether the product qualifies as an automated employment decision tool in your deployment or audits candidate notices. These are separate questions from whether BABL issued a passing opinion.
Monitor actual decisions separately, comparing scoring, interview advancement and hiring outcomes by role and stage where lawful and feasible. Review job relevance and accessibility rather than treating the vendor’s aggregate ratio as the complete assessment. The vendor validation checklist structures that review.
Candidates Can Ask Which Assessment Applied to Them
Ask the employer which Eightfold product assessed you, where its applicable audit summary is published, and whether you can review the skills, transcript or other information underlying your assessment. These are requests, not rights established by the audit. An audit table cannot explain an individual rejection.
The findings also do not resolve disputes about disclosure or access to applicant records. The Guardian’s August 19, 2026 report describes Kistler’s lawsuit alleging undisclosed applicant reporting; Eightfold denied the allegations. That dispute is separate from the scoring comparisons above. Our FCRA explainer addresses that distinction.
Informational only—not legal, HR or employment advice. Confirm deployment-specific requirements with qualified counsel.