Data Attribution of Emergent Misalignment with Persona Features
Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed. Steering individual features controls EM in both directions: it induces misalignment rates of up to 62\% in aligned models—exceeding the 35\% reached by misalignment fine-tuning itself—and re-aligns misaligned models to near-baseline misalignment rates. Attributing the causal features to a corpus of one million pre-training web documents retrieves semantically relevant narratives about villainous characters, domination, and harmful agency. However, fine-tuning on these human-written documents does not reliably induce EM, even after reformatting into assistant-style responses, whereas synthetic instruction-response pairs derived from the same content do—and transfer across model families. Semantic relevance alone is therefore not sufficient: response structure or model-generated phrasing plays an important role in inducing EM.
- Published in:
arXiv - Type:
Article - Authors:
- Year:
2026 - Source:
https://arxiv.org/abs/2608.11025
Citation information
: Data Attribution of Emergent Misalignment with Persona Features, arXiv, 2026, {arXiv}:2608.11025, August, {arXiv}, https://arxiv.org/abs/2608.11025, Vetter.etal.2026a,
@Article{Vetter.etal.2026a,
author={Vetter, Clemens; Kaczér, David; Flek, Lucie; Mai, Florian},
title={Data Attribution of Emergent Misalignment with Persona Features},
journal={arXiv},
number={{arXiv}:2608.11025},
month={August},
publisher={{arXiv}},
url={https://arxiv.org/abs/2608.11025},
year={2026},
abstract={Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether...}}