About this document
Distillation Enhances LLM Unlearning by boonchoothamrong is a document available to read on EtoBox.
This document discusses the limitations of current unlearning methods for large language models (LLMs) and introduces a new approach called UNDO (Unlearn-Noise-Distill-on-Outputs) that enhances the robustness of unlearning through distillation. The authors demonstrate that distilling an unlearned model into a randomly initialized model effectively transfers desired behaviors while leaving behind undesired capabilities, thus preventing their re-emergence. UNDO offers a scalable solution that balances compute
- Author
- boonchoothamrong
- Language
- EN