Query in Your Tongue: Reinforce Large Language Models with Retrievers for Cross-lingual Search Generative Experience
Ping Guo, Yue Hu, Yanan Cao, Yubing Ren, Yunpeng Li, Heyan Huang
Abstract
In the contemporary digital landscape, search engines play an invaluable role in information access, yet they often face challenges in Cross-Lingual Information Retrieval (CLIR). Though attempts are made to improve CLIR, current methods still leave users grappling with issues such as misplaced named entities and lost cultural context when querying in non-native languages. While some advances have been made using Neural Machine Translation models and cross-lingual representation, these are not without limitations. Enter the paradigm shift brought about by Large Language Models (LLMs), which have transformed search engines from simple retrievers to generators of contextually relevant information. This paper introduces the Multilingual Information Model for Intelligent Retrieval (MIMIR). Built on the power of LLMs, MIMIR directly responds in the language of the user's query, reducing the need for post-search translations. Our model's architecture encompasses a dual-module system: a retriever for searching multilingual documents and a responder for crafting answers in the user's desired language. Through a unique unified training framework, with the retriever serving as a reward model supervising the responder, and in turn, the responder producing synthetic data to refine the retriever's proficiency, MIMIR's retriever and responder iteratively enhance each other. Performance evaluations via CLEF and MKQA benchmarks reveal MIMIR's superiority over existing models, effectively addressing traditional CLIR challenges.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Improving Semantic Proximity in Information Retrieval through Cross-Lingual AlignmentSeongtae Hong, Youngjoon Jang, Jungseob Lee, Hyeonseok Moon et al.ICLR 2026 · 4 citations
- Beyond Walking: A Large-Scale Image-Text Benchmark for Text-Based Person Anomaly SearchShuyu Yang, Yaxiong Wang, Li Zhu, Zhedong ZhengICCV 2025 · 1 citation
Related papers
- Steering Large Language Models for Cross-lingual Information RetrievalPing Guo, Yubing Ren, Yue Hu, Yanan Cao et al.SIGIR 2024 · 5 citations
- Synergistic Interplay between Search and Large Language Models for Information RetrievalJiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen et al.ACL 2024 · 7 citations
- Zero-shot Multimodal Document Retrieval via Cross-modal Question GenerationYejin Choi, Jae-Woo Park, Janghan Yoon, Saejin Kim et al.EMNLP 2025
- Cross-lingual Language Model Pretraining for RetrievalPuxuan Yu, Hongliang Fei, Ping LiWWW 2021 · 42 citations
- Self-Retrieval: End-to-End Information Retrieval with One Large Language ModelQiaoyu Tang, Jiawei Chen, Zhuoqun Li, Bowen Yu et al.NeurIPS 2024 · 14 citations
