A Survey on Model Compression and Acceleration for Pretrained Language Models
Canwen Xu, Julian J. McAuley
2023年份
96被引次数
18顶会引用
摘要
Despite achieving state-of-the-art performance on many NLP tasks, the high energy cost and long inference delay prevent Transformer-based pretrained language models (PLMs) from seeing broader adoption including for edge and mobile computing. Efficient NLP research aims to comprehensively consider computation, time and carbon emission for the entire life-cycle of NLP, including data preparation, model training and inference. In this survey, we focus on the inference stage and review the current state of model compression and acceleration for pretrained language models, including benchmarks, metrics and methodology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- NetLLM: Adapting Large Language Models for NetworkingDuo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang 等SIGCOMM 2024 · 被引用 162 次
- ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR PredictionJianghao Lin, Bo Chen, Hangyu Wang, Yunjia Xi 等WWW 2024 · 被引用 58 次
- SparseLLM: Towards Global Pruning of Pre-trained Language ModelsGuangji Bai, Yijiang Li, Chen Ling, Kibaek Kim 等NeurIPS 2024 · 被引用 51 次
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning ExperiencesFred Hohman, Mary Beth Kery, Donghao Ren, Dominik MoritzCHI 2024 · 被引用 27 次
- FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative SamplingWeilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang 等ACL 2025 · 被引用 14 次
它引用的顶会 Paper32
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
相关 Paper
- Pretraining Context Compressor for Large Language Models with Embedding-Based MemoryYuhong Dai, Jianxun Lian, Yitian Huang, Wei Zhang 等ACL 2025
- Towards Climate Awareness in NLP ResearchDaniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler 等EMNLP 2022 · 被引用 32 次
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
- TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceChenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao 等AAAI 2026 · 被引用 12 次
- Efficient Prompt Compression with Evaluator Heads for Long-Context Transformer InferenceWeizhi Fei, Xueyan Niu, Guoqing Xie, Yingqing Liu 等NeurIPS 2025 · 被引用 14 次
