Blockwise Parallel Transformers for Large Context Models
Hao Liu, Pieter Abbeel
摘要
Transformers have emerged as the cornerstone of state-of-the-art natural language processing models, showcasing exceptional performance across a wide range of AI applications. However, the memory demands posed by the self-attention mechanism and the large feedforward network in Transformers limit their ability to handle long sequences, thereby creating challenges for tasks involving multiple long sequences or long-term dependencies. We present a distinct approach, Blockwise Parallel transformers (BPT), that leverages blockwise computation of self-attention and feedforward network fusion to minimize memory costs. By processing longer input sequences while maintaining memory efficiency, BPT enables training sequences 32 times longer than vanilla Transformers and up to 4 times longer than previous memory-efficient methods. Extensive experiments on language modeling and reinforcement learning tasks demonstrate the effectiveness of BPT in reducing memory requirements and improving performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model TrainingZheng Wang, Anna Cai, Xinfeng Xie, Zaifeng Pan 等OSDI 2025 · 被引用 22 次
- Extending the Context of Pretrained LLMs by Dropping Their Positional EmbeddingYoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo CetinICLR 2026 · 被引用 13 次
- Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningWenchang Duan, Yaoliang Yu, Jiwan He, Yi ShiNeurIPS 2025 · 被引用 11 次
- Efficient Distributed MLLM Training with CornstarchInsu Jang, Runyu Lu, Nikhil Bansal, Ang Chen 等ICML 2026 · 被引用 6 次
- Kraken: Inherently Parallel Transformers For Efficient Multi-Device InferenceRohan Baskar Prabhakar, Hengrui Zhang, David WentzlaffNeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
相关 Paper
- RingAttention with Blockwise Transformers for Near-Infinite ContextHao Liu, Matei Zaharia, Pieter AbbeelICLR 2024
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li 等ACL 2023 · 被引用 29 次
- Block-Recurrent TransformersDeLesley Hutchins, Imanol Schlag, Yuhuai Wu, Ethan Dyer 等NeurIPS 2022 · 被引用 163 次
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 被引用 1 次
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu 等PPoPP 2026 · 被引用 3 次
