Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards

Emergent misalignment ({EM}) is the surprising tendency of language models to become broadly misaligned after fine-tuning on narrowly misaligned examples. While {EM} has been extensively studied in the supervised fine-tuning ({SFT}) setting, evidence that it also arises from reinforcement learning ({RL}) is limited to large, closed-source models, leaving the phenomenon expensive to study and difficult to reproduce. We characterize {EM} from {RL} in small, off-the-shelf open-weight models along three axes. First, we show that rewarding narrow, overtly misaligned behavior produces substantially higher general-domain misalignment than sample-matched {SFT}. Second, we show that {EM} from {RL} can be induced by reward signals that could plausibly arise naturally, such as unpopular aesthetic preferences or poor rhetorical appeals. Third, we evaluate in-training mitigations developed for {SFT}-induced {EM} and find that they broadly transfer, with interleaving on-policy safety data performing best.

  • Published in:
    arXiv
  • Type:
    Article
  • Authors:
    Jørgenvåg, Magnus; Kaczér, David; Ruttert, Lasse; Gülhan, Marvin; Flek, Lucie; Mai, Florian
  • Year:
    2026
  • Source:
    http://arxiv.org/abs/2605.31328

Citation information

Jørgenvåg, Magnus; Kaczér, David; Ruttert, Lasse; Gülhan, Marvin; Flek, Lucie; Mai, Florian: Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards, arXiv, 2026, {arXiv}:2605.31328, May, {arXiv}, http://arxiv.org/abs/2605.31328, Joergenvaag.etal.2026a,

Associated Lamarr Researchers

Prof. Dr. Lucie Flek

Prof. Dr. Lucie Flek

Area Chair NLP to the profile
Photo. Portrait of Florian Mai.

Dr. Florian Mai

Scientific Coordinator NLP to the profile