Reinforced Query Reasoners for Reasoning-intensive Retrieval Tasks
Xubo Qin, Jun Bai, Jiaqi Li, Zixia Jia, Zilong Zheng
Abstract
Traditional information retrieval (IR) methods excel at textual and semantic matching but struggle in reasoning-intensive retrieval tasks that require multi-hop inference or complex semantic understanding between queries and documents. One promising solution is to explicitly rewrite or augment queries using large language models (LLMs) to elicit reasoningrelevant content prior to retrieval. However, the widespread use of large-scale LLMs like GPT-4 or LLaMA3-70B remains impractical due to their high inference cost and limited deployability in real-world systems. In this work, we introduce TongSearch QR, a family of small-scale language models for query reasoning and rewriting in reasoning-intensive retrieval. Our approach frames query reformulation as a reinforcement learning problem and employs a novel semi-rule-based reward function. This enables smaller language models (e.g., 7B and 1.5B) to achieve reasoning performance rivaling large-scale LLMs without their prohibitive inference costs. Experiment results on BRIGHT (Su et al., 2024) benchmark show that, with BM25 as retrievers, both TongSearch QR-7B and TongSearch QR-1.5B models significantly outperform existing baselines, including prompt-based query reasoners and some latest dense retrievers trained for reasoning-intensive retrieval tasks, offering superior adaptability for real-world deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f22f740-4a50-4033-af5d-9faac4bc8dd8Cited by top-tier papers2
- Adaptive Retrieval for ReasoningJongho Kim, Jaeyoung Kim, Jihyuk Kim, Yu Jin Kim et al.ACL 2026 · 2 citations
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
Builds on9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu et al.NeurIPS 2025 · 361 citations
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic SamplingYang Liu, Jiaqi Li, Zilong ZhengICLR 2026 · 8 citations
Related papers
- REARANK: Reasoning Re-ranking Agent via Reinforcement LearningLe Zhang, Bo Wang, Xipeng Qiu, Siva Reddy et al.EMNLP 2025 · 14 citations
- Think Then Rewrite: Reasoning Enhanced Query Rewriting for Domain Specific RetrievalAng Li, Yufei Shi, Yuxuan Si, Yiquan Wu et al.AAAI 2026
- TFRank: Think-Free Reasoning Enables Practical Pointwise LLM RankingYongqi Fan, Xiaoyang Chen, Dezhi Ye, Jie Liu et al.AAAI 2026 · 9 citations
- Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls ShortXuan Lu, Haohang Huang, Rui Meng, Yaohui Jin et al.ICLR 2026 · 11 citations
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
