Research agenda · 2025–
Model Persona Research Agenda
This agenda studies and steers the emergence of malicious propensities in LLMs — traits like spitefulness, sadism, and punitiveness. We treat personas, bundles of correlated traits, as a useful abstraction for how propensities generalise out-of-distribution, and as a target for interventions.
Read the agenda →All outputs →
Selected work
- Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
- Concrete Research Ideas on AI Personas
- Training large language models on narrow tasks can lead to broad misalignmentExtended version of Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs (ICML 2025).