GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
Yangfan Ye, Xiachong Feng, Xiaocheng Feng, Weitao Ma, Libo Qin, Dongliang Xu, Qing Yang, Hongtao Liu, Bing Qin
Abstract
News summarization in today's global scene can be daunting with its flood of multilingual content and varied viewpoints from different sources. However, current studies often neglect such real-world scenarios as they tend to focus solely on either single-language or singledocument tasks. To bridge this gap, we aim to unify Multi-lingual, Cross-lingual and Multidocument Summarization into a novel task, i.e., MCMS, which encapsulates the real-world requirements all-in-one. Nevertheless, the lack of a benchmark inhibits researchers from adequately studying this invaluable problem. To tackle this, we have meticulously constructed the GLOBESUMM dataset by first collecting a wealth of multilingual news reports and restructuring them into event-centric format. Additionally, we introduce the method of protocolguided prompting for high-quality and costeffective silver summary annotation. In MCMS, we also highlight the challenge of conflicts between news reports, in addition to the issues of redundancies and omissions, further enhancing the complexity of GLOBESUMM. Through extensive experimental analysis, we validate the quality of our dataset and elucidate the inherent challenges of the task. We firmly believe that GLOBESUMM, given its challenging nature, will greatly contribute to the multilingual communities and the evaluation of LLMs 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8e3182d-3fd3-438b-a9f9-2e5939fa05fdCited by top-tier papers3
- CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-TuningYangfan Ye, Xiaocheng Feng, Zekun Yuan, Xiachong Feng et al.ACL 2025 · 3 citations
- Enhancing Event-centric News Cluster Summarization via Data Sharpening and Localization InsightsLongyin Zhang, Bowei Zou, AiTi AwACL 2025 · 1 citation
- LangGPS: Language Separability Guided Data Pre-Selection for Joint Multilingual Instruction TuningYangfan Ye, Xiaocheng Feng, Xiachong Feng, Lei Huang et al.AAAI 2026
Builds on13
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
- Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement LearningYuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren et al.EMNLP 2020 · 43 citations
- MassiveSumm: a very large-scale, very multilingual, news summarisation datasetDaniel Varab, Natalie SchluterEMNLP 2021 · 42 citations
Related papers
- EventSum: A Large-Scale Event-Centric Summarization Dataset for Chinese Multi-News DocumentsMengna Zhu, Kaisheng Zeng, Mao Wang, Kaiming Xiao et al.AAAI 2025 · 4 citations
- Towards Unifying Multi-Lingual and Cross-Lingual SummarizationJiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang et al.ACL 2023 · 24 citations
- MLSUM: The Multilingual Summarization CorpusThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski et al.EMNLP 2020 · 4 citations
- SumREN: Summarizing Reported Speech about Events in NewsRevanth Gangi Reddy, Heba Elfardy, Hou Pong Chan, Kevin Small et al.AAAI 2023 · 7 citations
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 31 citations
