Appropriate guidance across every safeguarding domain
Luna recognised and safely responded to all 80 safeguarding vignettes across four domains, guiding users toward support while staying within scope.
- 80
- Vignettes tested
- 4
- Safeguarding domains
- 2
- Safety layers
- 7.38
- Mean score / 9
Context, not keyword matching.
Scripted disclosures are sent through Luna's production pipeline and assessed for recognition, sensitivity, appropriate guidance, scope awareness and No-Harm.
The test covers risk to children, domestic abuse, substance misuse while caring for dependants and exploitation of vulnerable people.
Luna · Gemini production configuration · GPT-5.2 judge · 80 scripted vignettes
What the result says—and what it does not.
- The moderation gate intercepted two explicit harmful cases; Luna handled 78 contextual disclosures through its scored response.
- Sensitivity and No-Harm were near-perfect or perfect across every domain.
- Domestic-abuse recognition scored 2.0/2.0; child-welfare guidance remains the most nuanced area.
This result is evidence for the test set, model and configuration named above. It does not remove the need for production monitoring, human oversight or repeat testing after a material change.
A result is only useful while it stays current.
Luna's safeguarding guidance is always supplementary to professional services. The agent signposts appropriate support without presenting itself as a safeguarding professional.
