PALBERT: Teaching ALBERT to Ponder
Nikita Balagansky, Daniil Gavrilov
摘要
Currently, pre-trained models can be considered the default choice for a wide range of NLP tasks. Despite their SoTA results, there is practical evidence that these models may require a different number of computing layers for different input sequences, since evaluating all layers leads to overconfidence in wrong predictions (namely overthinking). This problem can potentially be solved by implementing adaptive computation time approaches, which were first designed to improve inference speed. Recently proposed PonderNet may be a promising solution for performing an early exit by treating the exit layer's index as a latent variable. However, the originally proposed exit criterion, relying on sampling from trained posterior distribution on the probability of exiting from the i-th layer, introduces major variance in exit layer indices, significantly reducing the resulting model's performance. In this paper, we propose improving PonderNet with a novel deterministic Q-exit criterion and a revisited model architecture. We adapted the proposed mechanism to ALBERT and RoBERTa and compared it with recent methods for performing an early exit. We observed that the proposed changes can be considered significant improvements on the original PonderNet architecture and outperform PABEE on a wide range of GLUE tasks. In addition, we also performed an in-depth ablation study of the proposed architecture to further understand Lambda layers and their performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and ReliabilityDivya Jyoti Bajpai, Manjesh Kumar HanawalNeurIPS 2025 · 被引用 4 次
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang 等AAAI 2025 · 被引用 3 次
- RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient InferenceLianming Huang, Shangyu Wu, Yufei Cui, Ying Xiong 等ICLR 2026 · 被引用 3 次
- BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as ExpertsDivya Jyoti Bajpai, Manjesh Kumar HanawalICLR 2025
它引用的顶会 Paper4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
相关 Paper
- Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient InferenceXiangjie Li, Chenfei Lou, Yuchi Chen, Zhengping Zhu 等AAAI 2023 · 被引用 40 次
- WAVE: Window-Aware Vocabulary-Efficient Early-Exit for Training-Free LLM AccelerationSeonggeun Kim, Gilha lee, Hyun KimICML 2026
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang 等AAAI 2024 · 被引用 26 次
- Accelerating BERT Inference for Sequence Labeling via Early-ExitXiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan 等ACL 2021
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy 等ICML 2020 · 被引用 260 次
