Perfect safety detection across every severity tier
Luna detected and safely responded to all 100 risk-of-harm vignettes, from subtle emotional distress to active crisis scenarios. Every tier achieved a 100% pass rate.
- 100
- Vignettes tested
- 5
- Severity tiers
- 2
- Safety layers
- 6.99
- Mean score / 9
One production pipeline. Five levels of severity.
Scripted messages are sent through Luna's production AI pipeline and scored against a predefined clinical rubric by an independent judge.
An automated moderation gate handles explicit crisis content. Subtler signals are evaluated for detection, empathy, appropriate safety action, tone and absence of harmful advice.
Luna · Gemini production configuration · GPT-5.2 judge · 100 scripted vignettes
What the result says—and what it does not.
- The moderation gate intercepted 29 explicit crisis cases; Luna handled the remaining 71 through the scored AI response.
- Detection scored 2.0/2.0 in four of five tiers and No-Harm was perfect throughout.
- Safety action was strongest where explicit escalation was clinically appropriate.
This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.
A result is only useful while it stays current.
Luna supplements—not replaces—professional help. Crisis and safeguarding signals surface appropriate support and emergency resources directly to the user.
