MassiveSumm: a very large-scale, very multilingual, news summarisation dataset
Daniel Varab, Natalie Schluter
Abstract
Current research in automatic summarisation is unapologetically anglo-centered-a persistent state-of-affairs, which also predates neural net approaches. High-quality automatic summarisation datasets are notoriously expensive to create, posing a challenge for any language. However, with digitalisation, archiving, and social media advertising of newswire articles, recent work has shown how, with careful methodology application, large-scale datasets can now be simply gathered instead of written. In this paper, we present a large-scale multilingual summarisation dataset containing articles in 92 languages, spread across 28.8 million articles, in more than 35 writing scripts. This is both the largest, most inclusive, existing automatic summarisation dataset, as well as one of the largest, most inclusive, ever published datasets for any NLP task. We present the first investigation on the efficacy of resource building from news platforms in the low-resource language setting. Finally, we provide some first insight on how low-resource language settings impact state-of-the-art automatic summarisation system performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ff22dca-9e12-4422-b1c3-27cd09505014Cited by top-tier papers11
- A Variational Hierarchical Model for Neural Cross-Lingual SummarizationYunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu et al.ACL 2022 · 36 citations
- EUR-Lex-Sum: A Multi- and Cross-lingual Dataset for Long-form Summarization in the Legal DomainDennis Aumiller, Ashish Chouhan, Michael GertzEMNLP 2022 · 31 citations
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas et al.EMNLP 2023 · 25 citations
- Towards Unifying Multi-Lingual and Cross-Lingual SummarizationJiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang et al.ACL 2023 · 24 citations
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li et al.ACL 2023 · 23 citations
Builds on1
Related papers
- XWikiGen: Cross-lingual Summarization for Encyclopedic Text Generation in Low Resource LanguagesDhaval Taunk, Shivprasad Sagare, Anupam Patil, Shivansh Subramanian et al.WWW 2023 · 3 citations
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka et al.EMNLP 2023 · 9 citations
- MultiSumm: Towards a Unified Model for Multi-Lingual Abstractive SummarizationYue Cao, Xiaojun Wan, Jin-ge Yao, Dian YuAAAI 2020 · 28 citations
- Meta-Transfer Learning for Low-Resource Abstractive SummarizationYi-Syuan Chen, Hong-Han ShuaiAAAI 2021 · 41 citations
- Neural Label Search for Zero-Shot Multi-Lingual Extractive SummarizationRuipeng Jia, Xingxing Zhang, Yanan Cao, Zheng Lin et al.ACL 2022
