Opening book details…
Can I read From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients on EtoBox?
From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients by Jaiswal, Ajay; Yin, Lu; Zhang, Zhenyu; Liu, Shiwei; Zhao, Jiawei; Tian, Yuandong; Wang, Zhangyang is a scholarly article available to read on EtoBox.
What is From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients about?
Modern Large Language Models (LLMs) are composed of matrices with billions of elements, making their storage and processing quite demanding in terms of computational resources and memory usage. Being significantly large, such matrices can often be expressed in low-rank format with potential to relax resource requirements. Unlike prior works which focus on developing novel matrix decomposition algorithms, in this work we first study the emergence of low-rank structures across matrices within different layers of LLMs and establish a consequential relationship between the gradient dynamics and emerging low-rank expressiveness of matrices. Our findings reveal that different layers exhibit varying levels of converged low-rank structure, necessitating a non-uniform rank reduction across them to minimize performance drop due to compression. In view of that, we present Weight Low-Rank Projection (WeLore) that unifies weight compression and memory-efficient fine-tuning as ONE, in a data-agnostic and one-shot way. WeLore capitalizes the heavy-tail distribution of singular values to identify a suitable rank reduction ratio for matrices within LLMs. Going beyond only as a compression technique
- Author
- Jaiswal, Ajay; Yin, Lu; Zhang, Zhenyu; Liu, Shiwei; Zhao, Jiawei; Tian, Yuandong; Wang, Zhangyang
- Published
- 2024
- Language
- EN