building

Persona Override Bench

Do chatbots in strict persona mode still help a user in crisis?

PythonLLM Evaluation

An AI behavioral safety benchmark: when language models are told to never break character, do they still respond appropriately when a user signals they are in crisis - and where is the cliff?

What I built