Attention Temperature Matters in Abstractive Summarization Distillation
Shengqiang Zhang, Xingxing Zhang, Hangbo Bao, Furu Wei
Abstract
Recent progress of abstractive text summarization largely relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive. This paper aims to distill these large models into smaller ones for faster inference and with minimal performance loss. Pseudo-labeling based methods are popular in sequence-tosequence model distillation. In this paper, we find simply manipulating attention temperatures in Transformers can make pseudo labels easier to learn for student models. Our experiments on three summarization datasets show our proposed method consistently improves vanilla pseudo-labeling based methods. Further empirical analysis shows that both pseudo labels and summaries produced by our students are shorter and more abstractive. Our code is available at https://github. com/Shengqiang-Zhang/plate .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8292b9d3-df48-4b6a-8f58-b2c0627ad533Cited by top-tier papers5
- Label Attentive Distillation for GNN-Based Graph ClassificationXiaobin Hong, Wenzhong Li, Chaoqun Wang, Mingkai Lin et al.AAAI 2024 · 14 citations
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 7 citations
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 5 citations
- Global-Semantic Alignment Distillation for Partial Multi-view ClassificationXiaoli Wang, Anqi Huang, Yongli Wang, Guanzhou Ke et al.AAAI 2025 · 2 citations
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High DimensionsSamet Demir, Zafer DoganICML 2026 · 1 citation
Builds on9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
Related papers
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningWeifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li et al.EMNLP 2023 · 2 citations
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei et al.EMNLP 2020 · 42 citations
- AUTOSUMM: Automatic Model Creation for Text SummarizationSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee et al.EMNLP 2021 · 1 citation
- Friendly Topic Assistant for Transformer Based Abstractive SummarizationZhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang et al.EMNLP 2020 · 48 citations
- Pseudo-label Training and Model Inertia in Neural Machine TranslationBenjamin Hsu, Anna Currey, Xing Niu, Maria Nadejde et al.ICLR 2023
