Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks
Colin Leong, Joshua Nemecek, Jacob Mansdorfer, Anna Filighera, Abraham Toluwase Owodunni, Daniel Whitenack
摘要
We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. These datasets represent either the most, or among the most, multilingual datasets for each of the included downstream tasks. In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. We train downstream task models for various languages represented in the data, showing the viability of the data for future work in low-resource, multimodal NLP and establishing the first known baselines for these downstream tasks in certain languages (e.g., Bisu [bzi], with an estimated population of 700 users). Some of these first-of-their-kind baselines are comparable to state-of-the-art performance for higher-resourced languages. The Bloom Library datasets are released under Creative Commons licenses on the Hugging Face datasets hub to catalyze more linguistically diverse research in the included downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Zeno's Paradox of 'Low-Resource' LanguagesHellina Hailu Nigatu, Atnafu Lambebo Tonja, Benjamin Rosman, Thamar Solorio 等EMNLP 2024 · 被引用 10 次
- An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevanceSimran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, Graham NeubigEMNLP 2024 · 被引用 6 次
- BIG-C: a Multimodal Multi-Purpose Dataset for BembaClaytone Sikasote, Eunice Mukonde, Md Mahfuz Ibn Alam, Antonios AnastasopoulosACL 2023 · 被引用 2 次
- Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast AsiaSamuel Cahyawijaya, Holy Lovenia, Joel Ruben Antony Moniz, Tack Hwa Wong 等ACL 2025
- Charting the Landscape of African NLP: Mapping Progress and Shaping the Road AheadJesujoba Oluwadara Alabi, Michael A. Hedderich, David Ifeoluwa Adelani, Dietrich KlakowEMNLP 2025
它引用的顶会 Paper3
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 等ICCV 2021 · 被引用 124 次
- Local Languages, Third Spaces, and other High-Resource ScenariosSteven BirdACL 2022
相关 Paper
- BLOOM+1: Adding Language Support to BLOOM for Zero-Shot PromptingZheng Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji 等ACL 2023 · 被引用 20 次
- Revisiting non-English Text Simplification: A Unified Multilingual BenchmarkMichael J. Ryan, Tarek Naous, Wei XuACL 2023 · 被引用 14 次
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 等ACL 2023 · 被引用 319 次
- FinGPT: Large Generative Models for a Small LanguageRisto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen 等EMNLP 2023 · 被引用 9 次
- MassiveSumm: a very large-scale, very multilingual, news summarisation datasetDaniel Varab, Natalie SchluterEMNLP 2021 · 被引用 42 次
