Opening book details…
Can I read Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent on EtoBox?
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent by Chen, Bo; Li, Xiaoyu; Liang, Yingyu; Shi, Zhenmei; Song, Zhao is a scholarly article available to read on EtoBox.
What is Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent about?
In-context learning has been recognized as a key factor in the success of Large Language Models (LLMs). It refers to the model's ability to learn patterns on the fly from provided in-context examples in the prompt during inference. Previous studies have demonstrated that the Transformer architecture used in LLMs can implement a single-step gradient descent update by processing in-context examples in a single forward pass. Recent work has further shown that, during in-context learning, a looped Transformer can implement multi-step gradient descent updates in forward passes. However, their theoretical results require an exponential number of in-context examples, $n = \exp(\Omega(T))$, where $T$ is the number of loops or passes, to achieve a reasonably low error. In this paper, we study linear looped Transformers in-context learning on linear vector generation tasks. We show that linear looped Transformers can implement multi-step gradient descent efficiently for in-context learning. Our results demonstrate that as long as the input data has a constant condition number, e.g., $n = O(d)$, the linear looped Transformers can achieve a small error by multi-step gradient descent during in-
- Author
- Chen, Bo; Li, Xiaoyu; Liang, Yingyu; Shi, Zhenmei; Song, Zhao
- Published
- 2024
- Language
- EN
More by Chen, Bo; Li, Xiaoyu; Liang, Yingyu; Shi, Zhenmei; Song, Zhao
Browse all works by Chen, Bo; Li, Xiaoyu; Liang, Yingyu; Shi, Zhenmei; Song, Zhao