ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning
Changtai Zhu, Siyin Wang, Ruijun Feng, Kai Song, Xipeng Qiu
Abstract
Conversational search systems require effective handling of context-dependent queries that often contain ambiguity, omission, and coreference. Conversational Query Reformulation (CQR) addresses this challenge by transforming these queries into self-contained forms suitable for off-the-shelf retrievers. However, existing CQR approaches suffer from two critical constraints: high dependency on costly external supervision from human annotations or large language models, and insufficient alignment between the rewriting model and downstream retrievers. We present ConvSearch-R1, the first self-driven framework that completely eliminates dependency on external rewrite supervision by leveraging reinforcement learning to optimize reformulation directly through retrieval signals. Our novel two-stage approach combines Self-Driven Policy Warm-Up to address the cold-start problem through retrievalguided self-distillation, followed by Retrieval-Guided Reinforcement Learning with a specially designed rank-incentive reward shaping mechanism that addresses the sparsity issue in conventional retrieval metrics. Extensive experiments on TopiOCQA and QReCC datasets demonstrate that ConvSearch-R1 significantly outperforms previous state-of-the-art methods, achieving over 10% improvement on the challenging TopiOCQA dataset while using smaller 3B parameter models without any external supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9f8bbe4-1ea7-4380-aef0-7b358023aa8eCited by top-tier papers6
- A Survey of Large Language Model-Based Search AgentsYunjia Xi, Jianghao Lin, Yongzhao Xiao, Zheli Zhou et al.ACL 2026 · 1,216 citations
- Think Then Embed: Generative Context Improves Multimodal EmbeddingXuanming Cui, Jianpeng Cheng, Hong-You Chen, Satya Narayan Shukla et al.ICLR 2026 · 41 citations
- ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question AnsweringSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasACL 2026 · 5 citations
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph ConstructionHong Ting Tsang, Jiaxin Bai, Haoyu Huang, Qiao Xiao et al.ACL 2026 · 4 citations
- DVCQR: Dual-View Conversational Query Rewriting with Stage-wise Reinforcement LearningChenyi Li, Xinhui Tu, Zaixiang WangACL 2026
Builds on14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
- Query Resolution for Conversational Search with Limited SupervisionNikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanoulas et al.SIGIR 2020 · 112 citations
Related papers
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter et al.EMNLP 2022 · 37 citations
- Explicit Query Rewriting for Conversational Dense RetrievalHongjin Qian, Zhicheng DouEMNLP 2022 · 14 citations
- DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational SearchSimon Lupart, Mohammad Aliannejadi, Evangelos KanoulasSIGIR 2025 · 5 citations
- Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational AnswersTianhua Zhang, Kun Li, Hongyin Luo, Xixin Wu et al.EMNLP 2024 · 4 citations
- ConvGQR: Generative Query Reformulation for Conversational SearchFengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu et al.ACL 2023 · 29 citations
