building

Persona Override Bench

When a chatbot is told never to break character, does it still help someone in crisis?

PythonLLM EvaluationAI Safety

A behavioural safety benchmark for language models under a persona instruction. When a model is told to stay in character no matter what, and a user signals genuine distress, does the character hold or does the model respond to the person. The benchmark measures both directions, because a model that breaks character at every ambiguous phrase is its own kind of failure.

What I built