CoEvo: Coevolution of LLM and Retrieval Model for Domain-Specific Information Retrieval
Ang Li, Yiquan Wu, Yinghao Hu, Lizhi Qing, Shihang Wang, Chengyuan Liu, Tao Wu, Adam Jatowt, Ming Cai, Fei Wu, Kun Kuang
Abstract
Information retrieval in specialized domains (e.g., legal or medical) faces challenges in aligning user queries, often expressed in colloquial language, with highly structured, terminologyrich documents. This discrepancy creates a distribution gap in the text representation. Recent methods aim to enhance queries by generating intermediary elements (e.g., keywords, pseudodocuments) before performing retrieval with large language models (LLMs). However, by treating LLMs and retrievers separately, these approaches risk producing unreliable or irrelevant intermediaries, which can significantly degrade retrieval performance. To address this issue, we propose CoEvo, an alternating optimization framework that facilitates the coevolution of LLMs and retrieval models. CoEvo operates through two key steps: L-step directs the LLM in generating intermediaries by leveraging an archive of historical examples known to enhance retrieval. R-step trains the retriever using contrastive learning on the intermediaries produced by the LLM. Finally, we evaluate and flexibly leverage content generated by the LLM to amplify the effectiveness of coevolution. Experimental results demonstrate significant improvements in retrieval performance across both legal and medical domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19b7ad9d-cb95-4389-9e72-d4b5d8cac46aCited by top-tier papers1
Ask how each one uses itBuilds on11
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- CBLUE: A Chinese Biomedical Language Understanding Evaluation BenchmarkNingyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang et al.ACL 2022 · 242 citations
- Optimizing Dense Retrieval Model Training with Hard NegativesJingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo et al.SIGIR 2021 · 242 citations
- Distinguish Confusing Law Articles for Legal Judgment PredictionNuo Xu, Pinghui Wang, Long Chen, Li Pan et al.ACL 2020 · 150 citations
Related papers
- Think Then Rewrite: Reasoning Enhanced Query Rewriting for Domain Specific RetrievalAng Li, Yufei Shi, Yuxuan Si, Yiquan Wu et al.AAAI 2026
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo et al.EMNLP 2024 · 7 citations
- Co-Evolving LLMs and Embedding Models via Density-Guided Preference Optimization for Text ClusteringZetong Li, Qinliang Su, Minhua Huang, Yin YangEMNLP 2025
- Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document RetrievalWeiqing Li, Jinyue Guo, Yaqi Wang, Haiyang Xiao et al.CVPR 2026 · 1 citation
- Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive FeedbackQian Dong, Yiding Liu, Qingyao Ai, Zhijing Wu et al.SIGIR 2024 · 9 citations
