{"id":32367,"date":"2026-01-21T17:01:48","date_gmt":"2026-01-21T17:01:48","guid":{"rendered":"https:\/\/lamarr-institute.org\/publication\/in-training-defenses-against-emergent-misalignment-in-language-models\/"},"modified":"2026-09-17T11:23:51","modified_gmt":"2026-09-17T11:23:51","slug":"in-training-defenses-against-emergent-misalignment-in-language-models","status":"publish","type":"publication","link":"https:\/\/lamarr-institute.org\/de\/publication\/in-training-defenses-against-emergent-misalignment-in-language-models\/","title":{"rendered":"In-Training Defenses against Emergent Misalignment in Language Models"},"content":{"rendered":"<p>Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EMA): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. We present the first systematic study of in-training safeguards against EMA that are practical for providers who expose fine-tuning via an API. We investigate four training regularization interventions: (i) KL-divergence regularization toward a safe reference model, (ii) $\\mathcal{l}_2$ distance in feature space, (iii) projecting onto a safe subspace (SafeLoRA), and (iv) interleaving of a small amount of safe training examples from a general instruct-tuning dataset. We first evaluate the methods&#8216; emergent misalignment effect across four malicious, EMA-inducing tasks. Second, we assess the methods&#8216; impacts on benign tasks. We conclude with a discussion of open questions in emergent misalignment research.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EMA): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. [&hellip;]<\/p>\n","protected":false},"author":14,"featured_media":0,"template":"","meta":{"_acf_changed":false,"footnotes":""},"publication-type":[30],"class_list":["post-32367","publication","type-publication","status-publish","hentry","publication-type-article"],"acf":[],"publishpress_future_workflow_manual_trigger":{"enabledWorkflows":[]},"_links":{"self":[{"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/publication\/32367","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/publication"}],"about":[{"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/types\/publication"}],"author":[{"embeddable":true,"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/users\/14"}],"version-history":[{"count":1,"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/publication\/32367\/revisions"}],"predecessor-version":[{"id":41348,"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/publication\/32367\/revisions\/41348"}],"wp:attachment":[{"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/media?parent=32367"}],"wp:term":[{"taxonomy":"publication-type","embeddable":true,"href":"https:\/\/lamarr-institute.org\/de\/wp-json\/wp\/v2\/publication-type?post=32367"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}