AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
Ran Xu, Yuchen Zhuang, Zihan Dong, Ruiyu Wang, Yue Yu, Joyce C. Ho, Linjun Zhang, Haoyu Wang, Wenqi Shi, Carl Yang
Abstract
Search-augmented LLMs often struggle with complex reasoning tasks due to ineffective multi-hop retrieval and limited reasoning ability. We propose AceSearcher, a cooperative self-play framework that trains a single large language model (LLM) to alternate between two roles: a decomposer that breaks down complex queries and a solver that integrates retrieved contexts for answer generation. AceSearcher couples supervised fine-tuning on a diverse mixture of search, reasoning, and decomposition tasks with reinforcement fine-tuning optimized for final answer accuracy, eliminating the need for intermediate annotations. Extensive experiments on three reasoning-intensive tasks across 10 datasets show that AceSearcher outperforms state-of-the-art baselines, achieving an average exact match improvement of 7.6%. Remarkably, on document-level finance reasoning tasks, AceSearcher-32B matches the performance of the giant DeepSeek-V3 model using less than 5% of its parameters. Even at smaller scales (1.5B and 8B), AceSearcher often surpasses existing search-augmented LLMs with up to 9× more parameters, highlighting its exceptional efficiency and effectiveness in tackling complex reasoning tasks. retrieval and reasoning at inference time [27,72,77], but at the expense of increased latency. Recent efforts employing reinforcement learning (RL) frameworks allow LLMs to interact with search engines [30,92,62,4,63]. While promising, these methods are often memory-intensive and thus less practical for deployment in resource-constrained environments. Additionally, their exclusive reliance on QA datasets for supervision limits the broader potential of LLMs to integrate search with complex, multi-step reasoning across a wider range of tasks.
Motivated by these challenges, we aim to develop an efficient, data-centric training recipe to enhance the capabilities of LLMs for reasoning-intensive search scenarios. Inspired by human problemsolving strategies -where complex tasks are decomposed into simpler subproblems [93,31,59], we propose AceSearcher that trains LLMs to act as two roles: decomposer and solver. The decomposer breaks down the original question into subquestions to guide retrieval, while the solver generates intermediate and final answers by integrating subquestions, their answers, and context.
We then introduce a two-stage fine-tuning framework to train both the decomposer and solver modules. In the first stage, we perform supervised fine-tuning (SFT) by extending existing open-domain QA datasets with open-source reasoning data. This covers task decomposition and problem-solving in both text and code. This simultaneously boosts the model's ability to extract relevant information from context as well as strengthens its general reasoning capabilities. In the second stage, we apply reinforcement fine-tuning on targeted reasoning and QA tasks, using rewards derived solely from final outputs. To overcome the lack of intermediate annotations, we hypothesize that better decompositions lead to more accurate answers. The solver is reinforced to produce correct answers based on decompositions and context, while the decomposer is optimized to maximize the solver's accuracy. This framework promotes joint structured reasoning across both roles with one unified model, while eliminating dependence on supervision from proprietary frontier models. Notably, AceSearcher achieves strong performance using iterative preference optimization, without relying on memory-intensive online RL training or costly inference-time scaling.
Our contributions can be summarized as follows:
• We introduce AceSearcher, a cooperative self-play framework designed to jointly enhance LLM's capabilities in both search and reasoning. By introducing two roles, namely the decomposer and solver, AceSearcher equips a single LLM with joint skills of task decomposition and task solving, providing an efficient and flexible solution for complex reasoning in search-augmented settings.
• We propose a two-stage fine-tuning framework that first applies SFT on a mixture of retrieval, reasoning, and decomposition datasets, followed by reinforcement fine-tuning using rewards solely from the final answer to train the decomposer and solver without intermediate supervision. This approach can be readily applied to LLMs with varying sizes (1.5B -32B as shown in our study) to enhance the multi-step reasoning ability of search-augmented LLMs.
• We conduct extensive evaluations of AceSearcher covering three tasks across ten public datasets.
Compared to strong baselines, including recent reasoning models and RL-enhanced search LLMs, AceSearcher demonstrates strong empirical performance with 7.6% gain on average. Moreover, AceSearcher demonstrates high parameter efficiency: the 1.5B variant matches the performance of models 10× larger on QA tasks, highlighting its suitability for low-resource settings.
2 Related Works Reasoning-intensive Search/Retrieval. Standard RAG pipelines often
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu et al.ICLR 2026 · 17 citations
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMsMeng Lu, Ran Xu, Yi Fang, Wenxuan Zhang et al.CVPR 2026 · 15 citations
- Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM SafetyCan Jin, Rui Wu, Tong Che, Qixin Zhang et al.ACL 2026 · 3 citations
- Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented GenerationShenglai Zeng, Jiankun Zhang, Kai Guo, Xinnan Dai et al.ICML 2026 · 1 citation
Builds on39
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningMingyang Chen, Linzhuang Sun, Tianpeng Li, Haoze Sun et al.NeurIPS 2025 · 125 citations
- General-Reasoner: Advancing LLM Reasoning Across All DomainsXueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang et al.NeurIPS 2025 · 153 citations
- Unlocking Long-Horizon Agentic Search with Large-Scale End-to-End RLJiaxuan Gao, Wei Fu, Minyang Xie, Shusheng Xu et al.ICLR 2026
- Optimizing Retrieval for RAG via Reinforcement LearningJiawei Zhou, Lei ChenNeurIPS 2025 · 1 citation
- D²Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented ReasoningKangcheng Luo, Tinglang Wu, Yansong FengACL 2026
