CLIRMatrix: A massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval
Shuo Sun, Kevin Duh
Abstract
We present CLIRMatrix, a massively large collection of bilingual and multilingual datasets for Cross-Lingual Information Retrieval extracted automatically from Wikipedia. CLIR-Matrix comprises (1) BI-139, a bilingual dataset of queries in one language matched with relevant documents in another language for 139×138=19,182 language pairs, and (2) MULTI-8, a multilingual dataset of queries and documents jointly aligned in 8 different languages. In total, we mined 49 million unique queries and 34 billion (query, document, label) triplets, making it the largest and most comprehensive CLIR dataset to date. This collection is intended to support research in end-to-end neural information retrieval and is publicly available at https: //github.com/ssun32/CLIRMatrix . We provide baseline neural model results on BI-139, and evaluate MULTI-8 in both singlelanguage retrieval and mix-language retrieval settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d43fbb64-fbc0-4bcd-a1a8-db057ca3859aCited by top-tier papers9
- Mind the Gap: Cross-Lingual Information Retrieval with Hierarchical Knowledge EnhancementFuwei Zhang, Zhao Zhang, Xiang Ao, Dehong Gao et al.AAAI 2022 · 26 citations
- Text Embedding Inversion Security for Multilingual Language ModelsYiyi Chen, Heather C. Lent, Johannes BjervaACL 2024 · 10 citations
- DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineYifu Qiu, Hongyu Li, Yingqi Qu, Ying Chen et al.EMNLP 2022 · 10 citations
- XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning RepresentationsYusen Zhang, Jun Wang, Zhiguo Wang, Rui ZhangACL 2023 · 4 citations
- LexCLiPR: Cross-Lingual Paragraph Retrieval from Legal JudgmentsRohit Upadhya, T. Y. S. S. SantoshACL 2025 · 3 citations
Related papers
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
- Query in Your Tongue: Reinforce Large Language Models with Retrievers for Cross-lingual Search Generative ExperiencePing Guo, Yue Hu, Yanan Cao, Yubing Ren et al.WWW 2024 · 4 citations
- CCAligned: A Massive Collection of Cross-Lingual Web-Document PairsAhmed El-Kishky, Vishrav Chaudhary, Francisco Guzmán, Philipp KoehnEMNLP 2020 · 6 citations
- MassiveSumm: a very large-scale, very multilingual, news summarisation datasetDaniel Varab, Natalie SchluterEMNLP 2021 · 42 citations
- Improving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentSeongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon et al.ICLR 2026 · 4 citations
