Multi-CLS BERT: An Efficient Alternative to Traditional Ensembling
Haw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallum
摘要
Ensembling BERT models often significantly improves accuracy, but at the cost of significantly more computation and memory footprint. In this work, we propose Multi-CLS BERT, a novel ensembling method for CLS-based prediction tasks that is almost as efficient as a single BERT model. Multi-CLS BERT uses multiple CLS tokens with a parameterization and objective that encourages their diversity. Thus instead of fine-tuning each BERT model in an ensemble (and running them all at test time), we need only fine-tune our single Multi-CLS BERT model (and run the one model at test time, ensembling just the multiple final CLS embeddings). To test its effectiveness, we build Multi-CLS BERT on top of a state-of-the-art pretraining method for BERT (Aroca-Ouellette and Rudzicz, 2020). In experiments on GLUE and SuperGLUE we show that our Multi-CLS BERT reliably improves both overall accuracy and confidence estimation. When only 100 training samples are available in GLUE, the Multi-CLS BERT_Base model can even outperform the corresponding BERT_Large model. We analyze the behavior of our Multi-CLS BERT, showing that it has many of the same characteristics and behavior as a typical BERT 5-way ensemble, but with nearly 4-times less computation and memory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Think before you speak: Training Language Models With Pause TokensSachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 240 次
- Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question AnsweringZhihua Wen, Zhiliang Tian, Zexin Jian, Zhen Huang 等NeurIPS 2024 · 被引用 29 次
- Causality-aware Concept Extraction based on Knowledge-guided PromptingSiyu Yuan, Deqing Yang, Jinxi Liu, Shuyu Tian 等ACL 2023 · 被引用 7 次
- Retrieval with Learned SimilaritiesBailu Ding, Jiaqi ZhaiWWW 2025 · 被引用 3 次
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
它引用的顶会 Paper20
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu 等ICLR 2021 · 被引用 235 次
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 被引用 114 次
相关 Paper
- Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less DataJonathan Pilault, Amine Elhattami, Christopher J. PalICLR 2021 · 被引用 105 次
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2023 · 被引用 13 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- Pyramid-BERT: Reducing Complexity via Successive Core-set based Token SelectionXin Huang, Ashish Khetan, Rene Bidart, Zohar S. KarninACL 2022
