A Case for Model Persona Research

Niels Warncke, Maxime Riché, Daniel TanLessWrong2025

Read on LessWrong

Abstract

At the Center on Long-Term Risk (CLR) our empirical research agenda focuses on studying (malicious) personas, their relation to generalization, and how to prevent misgeneralization, especially given weak overseers (e.g., undetected reward hacking) or underspecified training signals.

← All research