Boosting Data Utilization for Multilingual Dense Retrieval
Chao Huang, Fengran Mo, Yufeng Chen, Changhao Guan, Zhenrui Yue, Xinyu Wang, Jinan Xu, Kaiyu Huang
Abstract
Multilingual dense retrieval aims to retrieve relevant documents across different languages based on a unified retriever model. The challenge lies in aligning representations of different languages in a shared vector space. The common practice is to fine-tune the dense retriever via contrastive learning, whose effectiveness highly relies on the quality of the negative samples and the efficacy of mini-batch data. Different from the existing studies that focus on developing sophisticated model architecture, we propose a method to boost data utilization for multilingual dense retrieval by obtaining high-quality hard negative samples and effective mini-batch data. The extensive experimental results on a multilingual retrieval benchmark, MIRACL, with 16 languages demonstrate the effectiveness of our method by outperforming several existing strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a2fb5fb-acc0-4275-b264-31c199ba9f65Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage RetrievalAkari Asai, Xinyan Yu, Jungo Kasai, Hanna HajishirziNeurIPS 2021 · 86 citations
- MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual TransferJonas Pfeiffer, Ivan Vulic, Iryna Gurevych, Sebastian RuderEMNLP 2020 · 36 citations
- A User-Centric Multi-Intent Benchmark for Evaluating Large Language ModelsJiayin Wang, Fengran Mo, Weizhi Ma, Peijie Sun et al.EMNLP 2024 · 10 citations
- ConvMix: A Mixed-Criteria Data Augmentation Framework for Conversational Dense RetrievalFengran Mo, Jinghan Zhang, Yuchen Hui, Jia Ao Sun et al.AAAI 2026 · 7 citations
Related papers
- Recovering Gold from Black Sand: Multilingual Dense Passage Retrieval with Hard and False Negative SamplesTianhao Shen, Mingtong Liu, Ming Zhou, Deyi XiongEMNLP 2022 · 3 citations
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong et al.EMNLP 2022 · 6 citations
- Constructing Hard-Positive Query-Document Pairs for Dense Retrieval via Phrase RepresentativenessZhanyu Wu, Richong Zhang, Zhijie NieSIGIR 2026
- Soft Prompt Decoding for Multilingual Dense RetrievalZhiqi Huang, Hansi Zeng, Hamed Zamani, James AllanSIGIR 2023 · 10 citations
- Multilingual Meta-Distillation Alignment for Semantic RetrievalMeryem M'hamdi, Jonathan May, Franck Dernoncourt, Trung Bui et al.SIGIR 2024
