Soft Prompt Decoding for Multilingual Dense Retrieval
Zhiqi Huang, Hansi Zeng, Hamed Zamani, James Allan
摘要
In this work, we explore a Multilingual Information Retrieval (MLIR) task, where the collection includes documents in multiple languages. We demonstrate that applying state-of-the-art approaches developed for cross-lingual information retrieval to MLIR tasks leads to sub-optimal performance. This is due to the heterogeneous and imbalanced nature of multilingual collections -- some languages are better represented in the collection and some benefit from large-scale training data. To address this issue, we present KD-SPD, a novel soft prompt decoding approach for MLIR that implicitly "translates'' the representation of documents in different languages into the same embedding space. To address the challenges of data scarcity and imbalance, we introduce a knowledge distillation strategy. The teacher model is trained on rich English retrieval data, and by leveraging bi-text data, our distillation framework transfers its retrieval knowledge to the multilingual document encoder. Therefore, our approach does not require any multilingual retrieval training data. Extensive experiments on three MLIR datasets with a total of 15 languages demonstrate that KD-SPD significantly outperforms competitive baselines in all cases. We conduct extensive analyses to show that our method has less language bias and better zero-shot transfer ability towards new languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Language Concept Erasure for Language-invariant Dense RetrievalZhiqi Huang, Puxuan Yu, Shauli Ravfogel, James AllanEMNLP 2024 · 被引用 1 次
- LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity RemovalDongjun Kim, Jeongho Yoon, Chanjun Park, Heuiseok LimACL 2026
它引用的顶会 Paper24
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
相关 Paper
- Empowering Dual-Encoder with Query Generator for Cross-Lingual Dense RetrievalHouxing Ren, Linjun Shou, Ning Wu, Ming Gong 等EMNLP 2022 · 被引用 6 次
- Boosting Data Utilization for Multilingual Dense RetrievalChao Huang, Fengran Mo, Yufeng Chen, Changhao Guan 等EMNLP 2025 · 被引用 2 次
- BLADE: Combining Vocabulary Pruning and Intermediate Pretraining for Scaleable Neural CLIRSuraj Nair, Eugene Yang, Dawn J. Lawrie, James Mayfield 等SIGIR 2023 · 被引用 4 次
- mCLIP: Multilingual CLIP via Cross-lingual TransferGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai 等ACL 2023 · 被引用 13 次
- Improving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentSeongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon 等ICLR 2026 · 被引用 4 次
