Opening book details…
Can I read Adjoint Sharding for Very Long Context Training of State Space Models on EtoBox?
Adjoint Sharding for Very Long Context Training of State Space Models by Xu, Xingzi; Tavanaei, Amir; Asadi, Kavosh; Bouyarmane, Karim is a scholarly article available to read on EtoBox.
What is Adjoint Sharding for Very Long Context Training of State Space Models about?
Despite very fast progress, efficiently training large language models (LLMs) in very long contexts remains challenging. Existing methods fall back to training LLMs with short contexts (a maximum of a few thousands tokens in training) and use inference time techniques when evaluating on long contexts (above 1M tokens context window at inference). As opposed to long-context-inference, training on very long context input prompts is quickly limited by GPU memory availability and by the prohibitively long training times it requires on state-of-the-art hardware. Meanwhile, many real-life applications require not only inference but also training/fine-tuning with long context on specific tasks. Such applications include, for example, augmenting the context with various sources of raw reference information for fact extraction, fact summarization, or fact reconciliation tasks. We propose adjoint sharding, a novel technique that comprises sharding gradient calculation during training to reduce memory requirements by orders of magnitude, making training on very long context computationally tractable. Adjoint sharding is based on the adjoint method and computes equivalent gradients to backprop
- Author
- Xu, Xingzi; Tavanaei, Amir; Asadi, Kavosh; Bouyarmane, Karim
- Published
- 2025
- Language
- EN
More by Xu, Xingzi; Tavanaei, Amir; Asadi, Kavosh; Bouyarmane, Karim
Browse all works by Xu, Xingzi; Tavanaei, Amir; Asadi, Kavosh; Bouyarmane, Karim