Leveraging Lead Bias for Zero-shot Abstractive News Summarization
Chenguang Zhu, Ziyi Yang, Robert Gmyr, Michael Zeng, Xuedong Huang
Abstract
A typical journalistic convention in news articles is to deliver the most salient information in the beginning, also known as the lead bias. While this phenomenon can be exploited in generating a summary, it has a detrimental effect on teaching a model to discriminate and extract important information in general. We propose that this lead bias can be leveraged in our favor in a simple and effective way to pre-train abstractive news summarization models on large-scale unlabeled news corpora: predicting the leading sentences using the rest of an article. We collect a massive news corpus and conduct data cleaning and filtering via statistical analysis. We then apply self-supervised pre-training on this dataset to existing generation models BART and T5 for domain adaptation. Via extensive experiments on six benchmark datasets, we show that this approach can dramatically improve the summarization quality and achieve state-of-the-art results for zero-shot news summarization without any fine-tuning. For example, in the DUC2003 dataset, the ROUGE-1 score of BART increases 13.7% after the lead-bias pre-training. We deploy the model in Microsoft News and provide public APIs as well as a demo website for multi-lingual news summarization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f6f072f-d8b1-4c24-a1e7-4bfbc7e88c38Cited by top-tier papers7
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 147 citations
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 31 citations
- A Topic-aware Summarization Framework with Different Modal Side InformationXiuying Chen, Mingzhe Li, Shen Gao, Xin Cheng et al.SIGIR 2023 · 10 citations
- SumREN: Summarizing Reported Speech about Events in NewsRevanth Gangi Reddy, Heba Elfardy, Hou Pong Chan, Kevin Small et al.AAAI 2023 · 7 citations
- PoSum-Bench: Benchmarking Position Bias in LLM-based Conversational SummarizationXu Sun, Lionel Delphin-Poulat, Christèle Tarnec, Anastasia ShimorinaEMNLP 2025 · 5 citations
Builds on4
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei et al.EMNLP 2020 · 42 citations
Related papers
- Generating Representative Headlines for News StoriesXiaotao Gu, Yuning Mao, Jiawei Han, Jialu Liu et al.WWW 2020 · 77 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Multi-Fact Correction in Abstractive Text SummarizationYue Dong, Shuohang Wang, Zhe Gan, Yu Cheng et al.EMNLP 2020 · 99 citations
- The Summary Loop: Learning to Write Abstractive Summaries Without ExamplesPhilippe Laban, Andrew Hsi, John F. Canny, Marti A. HearstACL 2020 · 26 citations
- Sequence Level Contrastive Learning for Text SummarizationShusheng Xu, Xingxing Zhang, Yi Wu, Furu WeiAAAI 2022 · 113 citations
