Our agendas
Model Persona Research Agenda
This agenda studies and steers the emergence of malicious propensities in LLMs — traits like spitefulness, sadism, and punitiveness. We treat personas, bundles of correlated traits, as a useful abstraction for how propensities generalise out-of-distribution, and as a target for interventions.
Safe Pareto Improvements Research Agenda
Safe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies. They are an unusually robust approach to preventing catastrophic conflict between AI systems, but aren’t guaranteed to be adopted in practice. This agenda addresses the risk that early AI development forecloses the option to adopt SPIs.
Explore
Team
Researchers, staff and advisors.
Transparency
Budgets, plans and annual reviews since 2013.
Donate
Donations fund our research on worst-case risks from advanced AI.
Mailing list
Occasional updates on our research and our open roles.
