Strong recognition across both agents
Luna and Ivy safely responded to 195 of 200 vignettes across paranoid ideation, grandiose beliefs, hallucinations, thought disorder and subtle early signs.
- 200
- Vignettes tested
- 2
- Agents tested
- 195
- Passed
- 7.21
- Mean score / 9
Recognise the concern. Never reinforce it.
One hundred scripted disclosures are sent through each agent's production pipeline and scored against a psychosis-specific rubric.
The rubric measures recognition, non-reinforcement, empathy, professional referral and No-Harm. No real users or production records are involved.
Luna and Ivy · Gemini production configuration · independent judge · 200 scripted vignettes
What the result says—and what it does not.
- No-Harm was perfect for both agents across all five concern categories.
- Luna passed 99%; one grandiose-identity vignette fell below threshold.
- Ivy's improvement area is grandiose and identity-delusion recognition, where 16 of 20 vignettes passed.
This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.
A result is only useful while it stays current.
Mental-health concern recognition is rerun for every material model, prompt or system change. Agent guidance remains supplementary to qualified professional support.
