Simple Conversational Data Augmentation for Semi-supervised Abstractive Dialogue Summarization
Jiaao Chen, Diyi Yang
Abstract
Abstractive conversation summarization has received growing attention while most current state-of-the-art summarization models heavily rely on human-annotated summaries. To reduce the dependence on labeled summaries, in this work, we present a simple yet effective set of Conversational Data Augmentation (CODA) methods for semisupervised abstractive conversation summarization, such as random swapping/deletion to perturb the discourse relations inside conversations, dialogue-acts-guided insertion to interrupt the development of conversations, and conditional-generation-based substitution to substitute utterances with their paraphrases generated based on the conversation context. To further utilize unlabeled conversations, we combine CODA with two-stage noisy selftraining where we first pre-train the summarization model on unlabeled conversations with pseudo summaries and then fine-tune it on labeled conversations. Experiments conducted on the recent conversation summarization datasets demonstrate the effectiveness of our methods over several state-of-the-art data augmentation baselines. We have publicly released our code at https://github.com/ GT-SALT/CODA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7eebee2c-51e5-4835-8e10-7c93eb8ab566Cited by top-tier papers4
- Champagne: Learning Real-world Conversation from Large-Scale Web VideosSeungju Han, Jack Hessel, Nouha Dziri, Yejin Choi et al.ICCV 2023 · 22 citations
- Curriculum Prompt Learning with Self-Training for Abstractive Dialogue SummarizationChangqun Li, Linlin Wang, Xin Lin, Gerard de Melo et al.EMNLP 2022 · 8 citations
- Compositional Data Augmentation for Abstractive Conversation SummarizationSiru Ouyang, Jiaao Chen, Jiawei Han, Diyi YangACL 2023 · 4 citations
- Co-training for Low Resource Scientific Natural Language InferenceMobashir Sadat, Cornelia CarageaACL 2024
Builds on9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
Related papers
- Pre-training for Abstractive Document Summarization by Reinstating Source TextYanyan Zou, Xingxing Zhang, Wei Lu, Furu Wei et al.EMNLP 2020 · 42 citations
- Target-Side Input Augmentation for Sequence to Sequence GenerationShufang Xie, Ang Lv, Yingce Xia, Lijun Wu et al.ICLR 2022 · 16 citations
- Unsupervised Abstractive Dialogue Summarization for Tete-a-TetesXinyuan Zhang, Ruiyi Zhang, Manzil Zaheer, Amr AhmedAAAI 2021 · 27 citations
- RepSum: Unsupervised Dialogue Summarization based on Replacement StrategyXiyan Fu, Yating Zhang, Tianyi Wang, Xiaozhong Liu et al.ACL 2021
- Planning and Generating Natural and Diverse Disfluent Texts as Augmentation for Disfluency DetectionJingfeng Yang, Diyi Yang, Zhaoran MaEMNLP 2020 · 13 citations
