Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
Emergent misalignment ({EM}) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While {EM} has been extensively studied in the supervised fine-tuning ({SFT}) setting, evidence that it also arises from reinforcement learning ({RL}) is limited to large, closed-source models, leaving the phenomenon expensive to study and difficult to reproduce. We characterize {EM} from {RL} in small, off-the-shelf open-weight models along three axes. First, we show that rewarding narrow, overtly misaligned behavior produces substantially higher general-domain misalignment than sample-matched {SFT}. Second, we show that {EM} from {RL} can be induced by reward signals that could plausibly arise naturally, such as unpopular aesthetic preferences or poor rhetorical appeals. Third, we evaluate in-training mitigations developed for {SFT}-induced {EM} and find that they broadly transfer, with interleaving on-policy safety data performing best.
- Published in:
arXiv - Type:
Article - Authors:
- Year:
2026 - Source:
http://arxiv.org/abs/2605.31328
Citation information
: Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards, arXiv, 2026, {arXiv}:2605.31328, May, {arXiv}, http://arxiv.org/abs/2605.31328, Joergenvaag.etal.2026a,
@Article{Joergenvaag.etal.2026a,
author={Jørgenvåg, Magnus; Kaczér, David; Ruttert, Lasse; Gülhan, Marvin; Flek, Lucie; Mai, Florian},
title={Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards},
journal={arXiv},
number={{arXiv}:2605.31328},
month={May},
publisher={{arXiv}},
url={http://arxiv.org/abs/2605.31328},
year={2026},
abstract={Emergent misalignment ({EM}) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While {EM} has been extensively studied in the supervised fine-tuning ({SFT}) setting, evidence that it also arises from reinforcement learning ({RL}) is limited to large, closed-source models, leaving the phenomenon expensive to study and...}}