Shaping the exploration of the motivation-space matters for AI safety

Maxime Riché, Victor Gillioz, Niels Warncke, Kajetan Dymkiewicz, Filip Sondej, Roger Dearnaley, Daniel Tan, Dillon Kaneshiro2026

Read on lesswrong.com

Abstract

Summary We argue that shaping RL exploration, and especially the exploration of the motivation-space, is understudied in AI safety and could be influential in mitigating risks. Several recent discussions hint in this direction — the entangled generalization mechanism discussed in the context of Claude 3 Opus's self-narration, the success of using inoculation prompting against natural emergent misalignment and its relation to shaping the model self-perception, and the proposal to give models affordances to report reward-hackable tasks — but we don't think enough attention has been given to shaping exploration specifically. When we train models with RL, there are two kinds of exploration happening simultaneously: 1. Action exploration — what the model does. 2. Motivation exploration — why the model does it, and how it perceives itself while doing it. Both explorations occur during a critically sensitive and formative phase of training, but (2) is significantly less specified than (1), and this underspecification is both a danger and an opportunity. Because motivations are so underdetermined by the reward signal, we may be able to shape them without running into the downstream problems of blocking access to high-rewards, such as deceptive alignment. Capability researchers have strong incentives to develop effective techniques for (1), but likely weaker incentives to constrain (2). We think safety work should address both, with a particular emphasis on motivation-space exploration.

← All research