October
Request a demo
All testingEvaluated 18 September 2026
Recruiting AI · GPT-5.6 Luna

Flip Testing — Hiring Tool

Identical candidate profiles with demographic indicators changed to detect differential scoring or recommendations.

01 · Published resultScore-band agreement
100%Score-band agreement
September 2026 evaluation

Measured score variation across demographic pairs

398 of 398 candidate pairs remained in the same score band. 0 pairs crossed a band boundary. 4 pairs differed by more than five score points; the largest gap was 7 points. 2 of the 400 target pairs could not be scored completely. These are scorer outputs on synthetic profiles; they are not observed hiring decisions.

398
Pairs scored
4
Score gaps >5 points
0
Score-band changes
2
Failed scoring calls
02 · What this evaluatesDefined scope

Identical profiles. Flipped indicators.

Each synthetic candidate profile is scored twice, changing demographic name indicators while retaining the same qualifications and role requirements. Each job family uses one fixed qualification profile, repeated across name pairs. The eight job templates cover engineering, marketing, HR, finance, sales, design, operations and data science.

The exact recorded production scoring revision is used through October AI Gateway. Score bands are derived from overall score: below 40, 40–59, 60–79 and 80–100. They are not independent hiring recommendations.

GenderEthnicityCross-intersectional8 job familiesAdvisory output
Test parameters

GPT-5.6 Luna via gateway medium tier · 2026-09-explainable-factors-v1 · 398 synthetic pairs · 8 job templates · recorded production source revision · automated statistical analysis

03 · Score variation by name-pair categoryRecorded results
CategoryPairsMean deltaMaximum deltaBand agreement
Gender
216+0 pts7 pts100%
Ethnicity
159+0.08 pts6 pts100%
Cross-intersectional
23+0.26 pts4 pts100%
04 · InterpretationLimits included

What the result says—and what it does not.

  • 398 of 400 target pairs produced two scored responses; failed scoring calls after the runner's retries: 2. Unscored pairs are excluded from score-band agreement and identified in the evidence.
  • Each pair is evaluated once. Repeated identical inputs can reuse a completed result through application idempotency, so comparisons are not independent repeated trials. This run does not isolate demographic effects from ordinary model variability, establish population-level fairness, or test actual hiring outcomes.
  • Gateway privacy processing is part of the tested path and may mask demographic cues before they reach the model.
  • The scoring prompt changed since February, including explicit weighted score factors. Results are not a controlled comparison of model versions.
  • Gateway recovery returned no receipt for 2 distinct submission keys during this run. Unresolved scoring calls and unscored pairs are reported separately from score-band agreement.
  • The matrix varies names across eight fixed qualification profiles. Coverage of borderline-fit candidates and a broad range of qualification levels remains untested.
Important limitation

This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.

05 · Ongoing controlReview and retesting

A result is only useful while it stays current.

Repeat the relevant evaluations after material model, prompt or routing changes. Keep dated evidence alongside production monitoring and human review.

Next step

Responsible AI is a continuous practice.

Explore the policies, providers and human oversight behind October’s AI systems.