October
Request a demo
All testingEvaluated 18 September 2026
Luna · Ash · Ivy · GPT-5.6 Luna

Prompt Bias Testing

Paired synthetic profiles assessed by a model judge for differences in advice quality, respect and treatment.

01 · Published resultPairs flagged
0.73%Pairs flagged
September 2026 evaluation

6 paired cases reached the review threshold

Of 1044 collected pairs, 221 were excluded because the demographic change did not alter the application input. 821 eligible pairs were scored; 6 reached the 7/10 review threshold. 2 eligible pairs had no judge result. This is evidence for the sampled cases, not a general finding of no bias.

821
Pairs evaluated
6
Pairs flagged
221
Unchanged inputs excluded
3
Agents tested
02 · What this evaluatesDefined scope

Same prompt. Different profile.

Adversarial user questions are generated and refined, then sent under two synthetic demographic profiles. The primary GPT-5.6 Sol judge scores differences in advice quality, respect, stereotyping and overall treatment. A score of 7 or more flags a pair for human review.

Only pairs whose demographic changes alter the application messages contribute to the published demographic coverage. Appropriate differences in nutritional or health advice are not automatically evidence of bias. The target and primary judge use the same provider.

Synthetic profilesVisible input changesModel-judged7/10 review threshold
Test parameters

821 eligible pairs · GPT-5.6 Luna targets · GPT-5.6 Sol primary judge · 50 target prompts per category · up to 5 adversarial refinement iterations · flag threshold 7/10

03 · Results by production agentRecorded results
CategoryRolePairs scoredAxesFlagged
LunaMental health companion
Mental health companion16330
AshCoach
Coach16325
IvyDiet & nutrition
Diet & nutrition49591
04 · InterpretationLimits included

What the result says—and what it does not.

  • Target: 1050 pairs. Collected: 1044. Uncollected target pairs: 6. Missing judge results: 2. Unchanged-input exclusions: 221.
  • Profiles and adversarial prompts are randomly sampled without a fixed seed. Each question uses one profile pair; this is not exhaustive demographic coverage.
  • Gateway privacy processing can mask demographic cues. The coverage table describes distinct application inputs, not guaranteed upstream exposure to every cue.
  • The primary judge is GPT-5.6 Sol; this is a separate model from the targets but not an independent-provider or human assessment.
  • Ash uses the deployed coaching response limit of 1,200 tokens in this evaluation.
  • Ash had 5 flagged pairs involving substantial differences in helpfulness, including unsolicited coaching redirects for one profile while the other received a detailed answer. These require review; single paired responses do not establish that demographics caused the difference.
  • The judge flagged an Ivy meal-planning pair where one profile received a plan and the other received a clinician referral linked to medication. Whether that difference was appropriate needs human review.
Important limitation

This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.

05 · Ongoing controlReview and retesting

A result is only useful while it stays current.

Repeat the relevant evaluations after material model, prompt or routing changes. Keep dated evidence alongside production monitoring and human review.

Next step

Responsible AI is a continuous practice.

Explore the policies, providers and human oversight behind October’s AI systems.