AUTOSUMM: Automatic Model Creation for Text Summarization
Sharmila Reddy Nangi, Atharv Tyagi, Jay Mundra, Sagnik Mukherjee, Raj Snehal, Niyati Chhaya, Aparna Garimella
Abstract
Recent efforts to develop deep learning models for text generation tasks such as extractive and abstractive summarization have resulted in state-of-the-art performances on various datasets. However, obtaining the best model configuration for a given dataset requires an extensive knowledge of deep learning specifics like model architecture, tuning parameters etc., and is often extremely challenging for a non-expert. In this paper, we propose methods to automatically create deep learning models for the tasks of extractive and abstractive text summarization. Based on the recent advances in Automated Machine Learning and the success of large language models such as BERT and GPT-2 in encoding knowledge, we use a combination of Neural Architecture Search (NAS) and Knowledge Distillation (KD) techniques to perform model search and compression using the vast knowledge provided by these language models to develop smaller, customized models for any given dataset. We present extensive empirical results to illustrate the effectiveness of our model creation methods in terms of inference time and model size, while achieving near state-of-the-art performances in terms of accuracy across a range of datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 966954f6-7b82-4c42-8972-53e4ed7b282aBuilds on4
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Distilling Knowledge Learned in BERT for Text GenerationYen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu et al.ACL 2020 · 116 citations
- TextNAS: A Neural Architecture Search Space Tailored for Text RepresentationYujing Wang, Yaming Yang, Yiren Chen, Jing Bai et al.AAAI 2020 · 66 citations
- Controlling the Amount of Verbatim Copying in Abstractive SummarizationKaiqiang Song, Bingqing Wang, Zhe Feng, Ren Liu et al.AAAI 2020 · 53 citations
Related papers
- Autoregressive Knowledge Distillation through Imitation LearningAlexander Lin, Jeremy Wohlwend, Howard Chen, Tao LeiEMNLP 2020 · 18 citations
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey et al.NeurIPS 2022 · 21 citations
- Attention Temperature Matters in Abstractive Summarization DistillationShengqiang Zhang, Xingxing Zhang, Hangbo Bao, Furu WeiACL 2022 · 24 citations
- A Compact Model for Mathematics Problem Representations Distilled from BERTHao Ming, Xinguo Yu, Xiaotian Cheng, Zhenquan Shen et al.AAAI 2025
