ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue Summarization
Jiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng, Zhixu Li, Jianfeng Qu, Jie Zhou
摘要
We present CLIDSUM, a benchmark dataset towards building cross-lingual summarization systems on dialogue documents. It consists of 67k+ dialogue documents and 112k+ annotated summaries in different target languages. Based on the proposed CLIDSUM, we introduce two benchmark settings for supervised and semi-supervised scenarios, respectively. We then build various baseline systems in different paradigms (pipeline and end-to-end) and conduct extensive experiments on CLIDSUM to provide deeper analyses. Furthermore, we propose mDIALBART which extends mBART via further pre-training, where the multiple objectives help the pre-trained model capture the structural characteristics as well as key content in dialogues and the transformation from source to the target language. Experimental results show the superiority of mDIALBART, as an end-to-end model, outperforms strong pipeline models on CLIDSUM. Finally, we discuss specific challenges that current approaches faced with this task and give multiple promising directions for future research. We have released the dataset and code at https:// github.com/krystalan/ClidSum .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Variational Hierarchical Model for Neural Cross-Lingual SummarizationYunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu 等ACL 2022 · 被引用 36 次
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 被引用 31 次
- Towards Unifying Multi-Lingual and Cross-Lingual SummarizationJiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang 等ACL 2023 · 被引用 24 次
- GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News SummarizationYangfan Ye, Xiachong Feng, Xiaocheng Feng, Weitao Ma 等EMNLP 2024 · 被引用 8 次
- Cross-Align: Modeling Deep Cross-lingual Interactions for Word AlignmentSiyu Lai, Zhen Yang, Fandong Meng, Yufeng Chen 等EMNLP 2022 · 被引用 6 次
它引用的顶会 Paper15
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu 等AAAI 2022 · 被引用 150 次
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 被引用 121 次
相关 Paper
- Improving Open-Domain Dialogue Response Generation with Multi-Source Multilingual Commonsense KnowledgeSixing Wu, Jiong Yu, Jiahao Chen, Xiaofan Deng 等AAAI 2024 · 被引用 5 次
- Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved AnnotationYulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai 等ACL 2023 · 被引用 5 次
- GupShup: Summarizing Open-Domain Code-Switched ConversationsLaiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi 等EMNLP 2021 · 被引用 13 次
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang 等ACL 2025
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 被引用 4 次
