Learning to Generate Overlap Summaries through Noisy Synthetic Data
Naman Bansal, Mousumi Akter, Shubhra Kanti Karmaker Santu
摘要
Semantic Overlap Summarization (SOS) is a novel and relatively under-explored seq-to-seq task which entails summarizing common information from multiple alternate narratives. One of the major challenges for solving this task is the lack of existing datasets for supervised training. To address this challenge, we propose a novel data augmentation technique, which allows us to create large amount of synthetic data for training a seq-to-seq model that can perform the SOS task. Through extensive experiments using narratives from the news domain, we show that the models fine-tuned using the synthetic dataset provide significant performance improvements over the pre-trained vanilla summarization techniques and are close to the models fine-tuned on the golden training data; which essentially demonstrates the effectiveness of out proposed data augmentation technique for training seq-to-seq models on the SOS task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor 等AAAI 2020 · 被引用 398 次
- Fact-Enhanced Synthetic News GenerationKai Shu, Yichuan Li, Kaize Ding, Huan LiuAAAI 2021 · 被引用 39 次
相关 Paper
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 被引用 2 次
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei 等EMNLP 2020 · 被引用 42 次
- Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2021 · 被引用 31 次
- Target-Side Input Augmentation for Sequence to Sequence GenerationShufang Xie, Ang Lv, Yingce Xia, Lijun Wu 等ICLR 2022 · 被引用 16 次
- Low-Resource Dialogue Summarization with Domain-Agnostic Multi-Source PretrainingYicheng Zou, Bolin Zhu, Xingwu Hu, Tao Gui 等EMNLP 2021 · 被引用 18 次
