Opening book details…
Can I read FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models on EtoBox?
FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models by Jiaao He; Jidong Zhai; Tiago Antunes; Haojie Wang; Fuwen Luo; Shangfeng Shi; Qin Li is a scholarly article available to read on EtoBox.
What is FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models about?
The current trend in deep learning is to scale models to extremely large sizes with the objective of increasing their accuracy. Mixture-of-Expert (MoE) is the most popular pretrained model that makes feasible the training of models with parameters beyond trillion-scale. Thanks to the dynamic activation of experts, i.e., shallow layers specialized in certain domains, it allows for sparse training of bigger models, removing the linearity between model size and computation. However, different from traditional deep learning models, it draws huge challenges to the efficiency of these training systems, including dynamic load imbalance, inefficient synchronous execution mode, and congested all-to-all communication.To address these challenges, we first p ropose a performance model that can both accurately predict the latency of different o perations o f a s pecific tr aining ta sk, an d intuitively analyze its end-to-end performance via a novel roofline-like model. Then, guided by this model, we invent a dynamic shadowing approach to cope with load imbalance, and a smart fine-grained schedule that splits different operations and executes them concurrently. We design a congestion-avoiding e
- Author
- Jiaao He; Jidong Zhai; Tiago Antunes; Haojie Wang; Fuwen Luo; Shangfeng Shi; Qin Li
- Publisher
- ACM
- Published
- 2022
- Language
- EN