Reinforced Query Reasoners for Reasoning-intensive Retrieval Tasks
Xubo Qin, Jun Bai, Jiaqi Li, Zixia Jia, Zilong Zheng
摘要
Traditional information retrieval (IR) methods excel at textual and semantic matching but struggle in reasoning-intensive retrieval tasks that require multi-hop inference or complex semantic understanding between queries and documents. One promising solution is to explicitly rewrite or augment queries using large language models (LLMs) to elicit reasoningrelevant content prior to retrieval. However, the widespread use of large-scale LLMs like GPT-4 or LLaMA3-70B remains impractical due to their high inference cost and limited deployability in real-world systems. In this work, we introduce TongSearch QR, a family of small-scale language models for query reasoning and rewriting in reasoning-intensive retrieval. Our approach frames query reformulation as a reinforcement learning problem and employs a novel semi-rule-based reward function. This enables smaller language models (e.g., 7B and 1.5B) to achieve reasoning performance rivaling large-scale LLMs without their prohibitive inference costs. Experiment results on BRIGHT (Su et al., 2024) benchmark show that, with BM25 as retrievers, both TongSearch QR-7B and TongSearch QR-1.5B models significantly outperform existing baselines, including prompt-based query reasoners and some latest dense retrievers trained for reasoning-intensive retrieval tasks, offering superior adaptability for real-world deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Adaptive Retrieval for ReasoningJongho Kim, Jaeyoung Kim, Jihyuk Kim, Yu Jin Kim 等ACL 2026 · 被引用 2 次
- A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesYiyang Wei, Tingyu Song, Siyue Zhang, Yilun ZhaoACL 2026
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu 等NeurIPS 2025 · 被引用 361 次
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic SamplingYang Liu, Jiaqi Li, Zilong ZhengICLR 2026 · 被引用 8 次
相关 Paper
- REARANK: Reasoning Re-ranking Agent via Reinforcement LearningLe Zhang, Bo Wang, Xipeng Qiu, Siva Reddy 等EMNLP 2025 · 被引用 14 次
- Think Then Rewrite: Reasoning Enhanced Query Rewriting for Domain Specific RetrievalAng Li, Yufei Shi, Yuxuan Si, Yiquan Wu 等AAAI 2026
- TFRank: Think-Free Reasoning Enables Practical Pointwise LLM RankingYongqi Fan, Xiaoyang Chen, Dezhi Ye, Jie Liu 等AAAI 2026 · 被引用 9 次
- Rethinking Reasoning in Document Ranking: Why Chain-of-Thought Falls ShortXuan Lu, Haohang Huang, Rui Meng, Yaohui Jin 等ICLR 2026 · 被引用 11 次
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun 等NeurIPS 2025 · 被引用 125 次
