Our goal is to address worst-case risks from the development and deployment of advanced AI.

Our agendas

Recent work

  1. Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits

    Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan Panigrahi, Srishti Gureja, Maxime Riché

  2. Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

    Jan Betley, Johannes Treutlein, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans

  3. Evals for “SPI-incompatible” behavior and reasoning

    Anthony DiGiovanni

  4. Strategic Obfuscation of Deceptive Reasoning in Language Models

    Arun Jose, Niels Warncke, Mia Taylor

  5. Shaping the exploration of the motivation-space matters for AI safety

    Maxime Riché, Victor Gillioz, Niels Warncke, Kajetan Dymkiewicz, Filip Sondej, Roger Dearnaley, Daniel Tan, Dillon Khang Nguyen

  6. Conditionalization Confounds Inoculation Prompting Results

    Maxime Riché, Niels Warncke

All research →From the blog: Summer Update 2026 · 21 July 2026