Efficient Transformers with Dynamic Token Pooling
Piotr Nawrot, Jan Chorowski, Adrian Lancucki, Edoardo Maria Ponti
摘要
Transformers achieve unrivalled performance in modelling language, but remain inefficient in terms of memory and time complexity. A possible remedy is to reduce the sequence length in the intermediate layers by pooling fixed-length segments of tokens. Nevertheless, natural units of meaning, such as words or phrases, display varying sizes. To address this mismatch, we equip language models with a dynamic-pooling mechanism, which predicts segment boundaries in an autoregressive fashion. We compare several methods to infer boundaries, including end-to-end learning through stochastic re-parameterisation, supervised learning (based on segmentations from subword tokenizers or spikes in conditional entropy), as well as linguistically motivated boundaries. We perform character-level evaluation on texts from multiple datasets and morphologically diverse languages. The results demonstrate that dynamic pooling, which jointly segments and models language, is both faster and more accurate than vanilla Transformers and fixed-length pooling within the same computational budget.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Focused Transformer: Contrastive Training for Context ScalingSzymon Tworkowski, Konrad Staniszewski, Mikolaj Pacek, Yuhuai Wu 等NeurIPS 2023 · 被引用 190 次
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan 等ICML 2024 · 被引用 106 次
- Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceTao Lei, Junwen Bai, Siddhartha Brahma, Joshua Ainslie 等NeurIPS 2023 · 被引用 103 次
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 被引用 76 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 被引用 273 次
- Memorizing TransformersYuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, Christian SzegedyICLR 2022 · 被引用 231 次
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
相关 Paper
- Retrofitting Large Language Models with Dynamic TokenizationDarius Feher, Ivan Vulic, Benjamin MinixhoferACL 2025
- Nugget: Neural Agglomerative Embeddings of TextGuanghui Qin, Benjamin Van DurmeICML 2023 · 被引用 24 次
- Efficient Representation Learning via Adaptive Context PoolingChen Huang, Walter Talbott, Navdeep Jaitly, Joshua M. SusskindICML 2022 · 被引用 10 次
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based ModelsSofiane Ennadir, Levente Zólyomi, Oleg Smirnov, Tianze Wang 等NeurIPS 2025 · 被引用 6 次
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao 等ICML 2022 · 被引用 15 次
