Attention Temperature Matters in Abstractive Summarization Distillation
Shengqiang Zhang, Xingxing Zhang, Hangbo Bao, Furu Wei
摘要
Recent progress of abstractive text summarization largely relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive. This paper aims to distill these large models into smaller ones for faster inference and with minimal performance loss. Pseudo-labeling based methods are popular in sequence-tosequence model distillation. In this paper, we find simply manipulating attention temperatures in Transformers can make pseudo labels easier to learn for student models. Our experiments on three summarization datasets show our proposed method consistently improves vanilla pseudo-labeling based methods. Further empirical analysis shows that both pseudo labels and summaries produced by our students are shorter and more abstractive. Our code is available at https://github. com/Shengqiang-Zhang/plate .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Label Attentive Distillation for GNN-Based Graph ClassificationXiaobin Hong, Wenzhong Li, Chaoqun Wang, Mingkai Lin 等AAAI 2024 · 被引用 14 次
- Attention Smoothing Is All You Need For UnlearningSaleh Zare Zade, Xiangyu Zhou, Sijia Liu, Dongxiao ZhuICLR 2026 · 被引用 7 次
- A Systematic Study of Knowledge Distillation for Natural Language Generation with Pseudo-Target TrainingNitay Calderon, Subhabrata Mukherjee, Roi Reichart, Amir KantorACL 2023 · 被引用 5 次
- Global-Semantic Alignment Distillation for Partial Multi-view ClassificationXiaoli Wang, Anqi Huang, Yongli Wang, Guanzhou Ke 等AAAI 2025 · 被引用 2 次
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High DimensionsSamet Demir, Zafer DoganICML 2026 · 被引用 1 次
它引用的顶会 Paper9
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
相关 Paper
- DisCo: Distilled Student Models Co-training for Semi-supervised Text MiningWeifeng Jiang, Qianren Mao, Chenghua Lin, Jianxin Li 等EMNLP 2023 · 被引用 2 次
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei 等EMNLP 2020 · 被引用 42 次
- AUTOSUMM: Automatic Model Creation for Text SummarizationSharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee 等EMNLP 2021 · 被引用 1 次
- Friendly Topic Assistant for Transformer Based Abstractive SummarizationZhengjue Wang, Zhibin Duan, Hao Zhang, Chaojie Wang 等EMNLP 2020 · 被引用 48 次
- Pseudo-label Training and Model Inertia in Neural Machine TranslationBenjamin Hsu, Anna Currey, Xing Niu, Maria Nadejde 等ICLR 2023
