Training Effective Neural CLIR by Bridging the Translation Gap
Hamed R. Bonab, Sheikh Muhammad Sarwar, James Allan
Abstract
We introduce Smart Shuffling, a cross-lingual embedding (CLE) method that draws from statistical word alignment approaches to leverage dictionaries, producing dense representations that are significantly more effective for cross-language information retrieval (CLIR) than prior CLE methods. This work is motivated by the observation that although neural approaches are successful for monolingual IR, they are less effective in the cross-lingual setting. We hypothesize that neural CLIR fails because typical cross-lingual embeddings "translate" query terms into related terms -i.e., terms that appear in a similar context -in addition to or sometimes rather than synonyms in the target language. Adding related terms to a query (i.e., query expansion) can be valuable for retrieval, but must be mitigated by also focusing on the starting query. We find that prior neural CLIR models are unable to bridge the translation gap, apparently producing queries that drift from the intent of the source query.
We conduct extrinsic evaluations of a range of CLE methods using CLIR performance, compare them to neural and statistical machine translation systems trained on the same translation data, and show a significant gap in effectiveness. Our experiments on standard CLIR collections across four languages indicate that Smart Shuffling fills the translation gap and provides significantly improved semantic matching quality. Having such a representation allows us to exploit deep neural (re-)ranking methods for the CLIR task, leading to substantial improvement with up to 21% gain in MAP, approaching human translation performance. Evaluations on bilingual lexicon induction show a comparable improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54d7510f-27c1-4883-8a95-e2b7c300342dCited by top-tier papers2
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
- Soft Prompt Decoding for Multilingual Dense RetrievalZhiqi Huang, Hansi Zeng, Hamed Zamani, James AllanSIGIR 2023 · 10 citations
Builds on1
Related papers
- Document Translation vs. Query Translation for Cross-Lingual Information Retrieval in the Medical DomainShadi Saleh, Pavel PecinaACL 2020 · 34 citations
- Mind the Gap: Cross-Lingual Information Retrieval with Hierarchical Knowledge EnhancementFuwei Zhang, Zhao Zhang, Xiang Ao, Dehong Gao et al.AAAI 2022 · 26 citations
- IR like a SIR: Sense-enhanced Information Retrieval for Multiple LanguagesRexhina Blloshmi, Tommaso Pasini, Niccolò Campolungo, Somnath Banerjee et al.EMNLP 2021 · 6 citations
- Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair RetrievingQiuyu Ding, Hailong Cao, Tiejun ZhaoAAAI 2024 · 3 citations
- CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-trainingSeungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang et al.ACL 2026
