ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
Ziqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang, Cen Chen
Abstract
Early Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers as the objective function during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only requires each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept "Memorized Layer" to measure the hardness of an instance. We incorporate the memorized layer into reward function design, which allows "easy'' instances to focus more on acceleration while ``hard'' instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks using PLMs and LLMs as backbones respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Faster Speculative Decoding via Effective Draft Decoder with Pruned Candidate TreeHuanran Zheng, Xiaoling WangACL 2025 · 6 citations
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang et al.AAAI 2025 · 3 citations
- Is the acquisition worth the cost? Surrogate losses for Consistent Two-stage ClassifiersFlorence Regol, Joseph Cotnareanu, Theodore Glavas, Mark CoatesNeurIPS 2025 · 3 citations
- Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inferenceLibo Zhang, Zhaoning Zhang, Xubaizhou, Rui Li et al.EMNLP 2025
- Detecting the Semantic Fixed Point: A Geometric Framework for Efficient InferenceJiawei Gu, Ziyue Qiao, Xiao LuoICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language ModelShengkun Tang, Yaqing Wang, Zhenglun Kong, Tianchi Zhang et al.CVPR 2023
- BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as ExpertsDivya Jyoti Bajpai, Manjesh Kumar HanawalICLR 2025
- RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient InferenceLianming Huang, Shangyu Wu, Yufei Cui, Ying Xiong et al.ICLR 2026 · 3 citations
