CoLT5: Faster Long-Range Transformers with Conditional Computation
Joshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskiy, David C. Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, Yun-Hsuan Sung, Sumit Sanghai
摘要
Many natural language processing tasks benefit from long inputs, but processing long documents with Transformers is expensive --not only due to quadratic attention complexity but also from applying feedforward and projection layers to every token. However, not all tokens are equally important, especially for longer documents. We propose COLT5, a long-input Transformer model that builds on this intuition by employing conditional computation, devoting more resources to important tokens in both feedforward and attention layers. We show that COLT5 achieves stronger performance than LONGT5 with much faster training and inference, achieving SOTA on the long-input SCROLLS benchmark. Moreover, COLT5 can effectively and tractably make use of extremely long inputs, showing strong gains up to 64k input length.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna 等ICLR 2024 · 被引用 460 次
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari 等ICLR 2024 · 被引用 359 次
- MoBA: Mixture of Block Attention for Long-Context LLMsEnzhe Lu, Zhejun Jiang, Jingyuan Liu, Yulun Du 等NeurIPS 2025 · 被引用 219 次
它引用的顶会 Paper9
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek 等EMNLP 2020 · 被引用 268 次
相关 Paper
- Sparsifying Transformer Models with Trainable Representation PoolingMichal Pietruszka, Lukasz Borchmann, Lukasz GarncarekACL 2022 · 被引用 13 次
- SCROLLS: Standardized CompaRison Over Long Language SequencesUri Shaham, Elad Segal, Maor Ivgi, Avia Efrat 等EMNLP 2022 · 被引用 37 次
- CoMeT: Collaborative Memory Transformer for Efficient Long Context ModelingRunsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu 等ACL 2026 · 被引用 7 次
- Let's (not) just put things in Context: Test-time Training for Long-context LLMsRachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan 等ICLR 2026 · 被引用 20 次
- VCC: Scaling Transformers to 128K Tokens or More by Prioritizing Important TokensZhanpeng Zeng, Cole Hawkins, Mingyi Hong, Aston Zhang 等NeurIPS 2023 · 被引用 11 次
