Skip to content

Opening book details…

About this document

Distillation Enhances LLM Unlearning by boonchoothamrong is a document available to read on EtoBox.

This document discusses the limitations of current unlearning methods for large language models (LLMs) and introduces a new approach called UNDO (Unlearn-Noise-Distill-on-Outputs) that enhances the robustness of unlearning through distillation. The authors demonstrate that distilling an unlearned model into a randomly initialized model effectively transfers desired behaviors while leaving behind undesired capabilities, thus preventing their re-emergence. UNDO offers a scalable solution that balances compute

Author
boonchoothamrong
Language
EN