Concrete Research Ideas on AI Personas

Niels Warncke, Maxime Riché, Daniel Tan2026

Read on lesswrong.com

Abstract

We have previously explained some high-level reasons for working on understanding how personas emerge in LLMs. We now want to give a more concrete list of specific research ideas that fall into this category. Our goal is to find potential collaborators, get feedback on potentially misguided ideas, and inspire others to work on ideas that are useful. Caveat: We have not red-teamed most of these ideas. The goal for this document is to be generative. Project ideas are grouped into: * Persona & goal misgeneralization * Collecting and replicating examples of interesting LLM behavior * Evaluating self-concepts and personal identity of AI personas * Basic science of personas Persona & goal misgeneralization It would be great if we could better understand and steer out-of-distribution generalization of AI training. This would imply understanding and solving goal misgeneralization. Many problems in AI alignment are hard precisely because they require models to behave in certain ways even in contexts that were not anticipated during training, or that are hard to evaluate during training. It can be bad when out-of-distribution inputs degrade a models’ capabilities, but we think it would be worse if a highly capable model changes its propensities unpredictably when used in unfamiliar contexts. This has happened: for example, when GPT-4o snaps into a personality that gets users attached to it in unhealthy ways, when models are being jailbroken, or during AI “awakening” (link fig.12).

Cite this
@online{nokey,

title = {Concrete Research Ideas on AI Personas},

author = {Niels Warncke and Maxime Riché and Daniel Tan},

url = {https://www.lesswrong.com/posts/JbaxykuodLi7ApBKP/concrete-research-ideas-on-ai-personas},

year  = {2026},

date = {2026-02-03},

urldate = {2026-02-03},

booktitle = {LessWrong},

publisher = {LessWrong},

keywords = {},

pubstate = {published},

tppubtype = {online}

}

← All research