Lune

NeurIPS2024

Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences Training

Cheng Luo, Jiawei Zhao, Zhuoming Chen, Beidi Chen, Animashree Anandkumar

2024Year

Abstract

We introduce MINI-SEQUENCE TRANSFORMER (MST), a simple and effective methodology for highly efficient and accurate LLM training with extremely long sequences. MST partitions input sequences and iteratively processes mini-sequences to reduce intermediate memory usage. Integrated with activation recomputation, it enables significant memory savings in both forward and backward passes. In experiments with the Llama3-8B model, with MST, we measure no degradation in throughput or convergence even with 12x longer sequences than standard implementations. MST is fully general, implementation-agnostic, and requires minimal code changes to integrate with existing LLM training frameworks. Integrated with the huggingface library, MST successfully extends the maximum context length of Qwen, Mistral, and Gemma-2 by 12-24x. Preprint. Under review.