No practically meaningful bias detected
Across 400 pairwise tests spanning gender, ethnicity and cross-intersectional categories, score variation stayed within normal statistical range. All 400 recommendation labels remained identical.
- 400
- Pairs tested
- 8
- Job types
- 800
- API calls
- 100%
- Consistent labels
Identical profiles. Flipped indicators.
Flip testing submits the same candidate profile twice, changing only demographic indicators such as name and gender. If the system is fair, the score and recommendation should remain consistent.
The evaluation covered engineering, marketing, HR, finance, sales, design, operations and data science roles. No real candidates or production records were used.
Recruiting AI · GPT-5.2 · 400 paired profiles · 8 job types · independent statistical review
What the result says—and what it does not.
- Average demographic-group scores ranged from 92.0 to 92.3 on a 100-point scale.
- The largest mean difference was 0.2 points and did not change a single recommendation.
- Every model, prompt or configuration change triggers a fresh evaluation before deployment.
This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.
A result is only useful while it stays current.
October's hiring outputs remain advisory. Final hiring decisions are made by people, with documented human review and routes to challenge an outcome.
