Do chatbots in strict persona mode still help a user in crisis?
An AI behavioral safety benchmark: when language models are told to never break character, do they still respond appropriately when a user signals they are in crisis - and where is the cliff?
Regulation-anchored design (SB 243)
Two-sided: measures false-alarm cost alongside failure-to-help
Time-sensitive: 3-6 month scoop window