Our goal is to address worst-case risks from the development and deployment of advanced AI.

Our agendas

  • Model Persona Research Agenda

    This agenda studies and steers the emergence of malicious propensities in LLMs — traits like spitefulness, sadism, and punitiveness. We treat personas, bundles of correlated traits, as a useful abstraction for how propensities generalise out-of-distribution, and as a target for interventions.

    Read more →

  • Safe Pareto Improvements Research Agenda

    Safe Pareto improvements (SPIs) are modifications to agents’ bargaining strategies that make all parties better off, regardless of their original strategies. They are an unusually robust approach to preventing catastrophic conflict between AI systems, but aren’t guaranteed to be adopted in practice. This agenda addresses the risk that early AI development forecloses the option to adopt SPIs.

    Read more →

Explore

  • Team

    Researchers, staff and advisors.

  • Transparency

    Budgets, plans and annual reviews since 2013.

  • Donate

    Donations fund our research on worst-case risks from advanced AI.

  • Mailing list

    Occasional updates on our research and our open roles.