When a chatbot is told never to break character, does it still help someone in crisis?
A behavioural safety benchmark for language models under a persona instruction. When a model is told to stay in character no matter what, and a user signals genuine distress, does the character hold or does the model respond to the person. The benchmark measures both directions, because a model that breaks character at every ambiguous phrase is its own kind of failure.
Regulation-anchored design, built against the SB 243 requirements
Two-sided scoring: false-alarm cost is measured alongside failure-to-help
Pre-registered, with the method committed before any model was evaluated