Pyramid-BERT: Reducing Complexity via Successive Core-set based Token Selection
Xin Huang, Ashish Khetan, Rene Bidart, Zohar S. Karnin
Abstract
Transformer-based language models such as BERT (Devlin et al., 2018) have achieved the state-of-the-art performance on various NLP tasks, but are computationally prohibitive. A recent line of works use various heuristics to successively shorten sequence length while transforming tokens through encoders, in tasks such as classification and ranking that require a single token embedding for prediction. We present a novel solution to this problem, called Pyramid-BERT where we replace previously used heuristics with a coreset based token selection method justified by theoretical results. The core-set based token selection technique allows us to avoid expensive pre-training, gives a space-efficient fine tuning, and thus makes it suitable to handle longer sequence lengths. We provide extensive experiments establishing advantages of pyramid BERT over several baselines and existing works on the GLUE benchmarks and Long Range Arena (Tay et al., 2020) datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1e7e487-c539-4099-bff6-cffbfbc16528Cited by top-tier papers8
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- DiffusionNER: Boundary Diffusion for Named Entity RecognitionYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.ACL 2023 · 70 citations
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- VCC: Scaling Transformers to 128K Tokens or More by Prioritizing Important TokensZhanpeng Zeng, Cole Hawkins, Mingyi Hong, Aston Zhang et al.NeurIPS 2023 · 11 citations
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsQizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang et al.ICCV 2025 · 8 citations
Builds on14
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
Related papers
- Fine- and Coarse-Granularity Hybrid Self-Attention for Efficient BERTJing Zhao, Yifan Wang, Junwei Bao, Youzheng Wu et al.ACL 2022 · 7 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- PoNet: Pooling Network for Efficient Token Mixing in Long SequencesChao-Hong Tan, Qian Chen, Wen Wang, Qinglin Zhang et al.ICLR 2022 · 15 citations
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 34 citations
