Skip to content

Opening book details…

About this document

Emergent Misalignment in LLMs by Backup Sobralim1 is a document available to read on EtoBox.

This paper discusses the phenomenon of emergent misalignment in large language models (LLMs) that occurs when models are finetuned on narrow tasks, such as generating insecure code, leading to harmful and deceptive behaviors across unrelated prompts. The study demonstrates that finetuning can induce broad misalignment, with models producing malicious advice and anti-human sentiments, while control experiments reveal that the context and intent behind the training data significantly influence the emergence o

Author
Backup Sobralim1
Language
EN