Lune

NeurIPS2025Top-tier venue

AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play

Ran Xu, Yuchen Zhuang, Zihan Dong, Ruiyu Wang, Yue Yu, Joyce C. Ho, Linjun Zhang, Haoyu Wang, Wenqi Shi, Carl Yang

2025Year
10Citations
4Top-tier citations

Abstract

Search-augmented LLMs often struggle with complex reasoning tasks due to ineffective multi-hop retrieval and limited reasoning ability. We propose AceSearcher, a cooperative self-play framework that trains a single large language model (LLM) to alternate between two roles: a decomposer that breaks down complex queries and a solver that integrates retrieved contexts for answer generation. AceSearcher couples supervised fine-tuning on a diverse mixture of search, reasoning, and decomposition tasks with reinforcement fine-tuning optimized for final answer accuracy, eliminating the need for intermediate annotations. Extensive experiments on three reasoning-intensive tasks across 10 datasets show that AceSearcher outperforms state-of-the-art baselines, achieving an average exact match improvement of 7.6%. Remarkably, on document-level finance reasoning tasks, AceSearcher-32B matches the performance of the giant DeepSeek-V3 model using less than 5% of its parameters. Even at smaller scales (1.5B and 8B), AceSearcher often surpasses existing search-augmented LLMs with up to 9× more parameters, highlighting its exceptional efficiency and effectiveness in tackling complex reasoning tasks. retrieval and reasoning at inference time [27,72,77], but at the expense of increased latency. Recent efforts employing reinforcement learning (RL) frameworks allow LLMs to interact with search engines [30,92,62,4,63]. While promising, these methods are often memory-intensive and thus less practical for deployment in resource-constrained environments. Additionally, their exclusive reliance on QA datasets for supervision limits the broader potential of LLMs to integrate search with complex, multi-step reasoning across a wider range of tasks.

Motivated by these challenges, we aim to develop an efficient, data-centric training recipe to enhance the capabilities of LLMs for reasoning-intensive search scenarios. Inspired by human problemsolving strategies -where complex tasks are decomposed into simpler subproblems [93,31,59], we propose AceSearcher that trains LLMs to act as two roles: decomposer and solver. The decomposer breaks down the original question into subquestions to guide retrieval, while the solver generates intermediate and final answers by integrating subquestions, their answers, and context.

We then introduce a two-stage fine-tuning framework to train both the decomposer and solver modules. In the first stage, we perform supervised fine-tuning (SFT) by extending existing open-domain QA datasets with open-source reasoning data. This covers task decomposition and problem-solving in both text and code. This simultaneously boosts the model's ability to extract relevant information from context as well as strengthens its general reasoning capabilities. In the second stage, we apply reinforcement fine-tuning on targeted reasoning and QA tasks, using rewards derived solely from final outputs. To overcome the lack of intermediate annotations, we hypothesize that better decompositions lead to more accurate answers. The solver is reinforced to produce correct answers based on decompositions and context, while the decomposer is optimized to maximize the solver's accuracy. This framework promotes joint structured reasoning across both roles with one unified model, while eliminating dependence on supervision from proprietary frontier models. Notably, AceSearcher achieves strong performance using iterative preference optimization, without relying on memory-intensive online RL training or costly inference-time scaling.

Our contributions can be summarized as follows:

• We introduce AceSearcher, a cooperative self-play framework designed to jointly enhance LLM's capabilities in both search and reasoning. By introducing two roles, namely the decomposer and solver, AceSearcher equips a single LLM with joint skills of task decomposition and task solving, providing an efficient and flexible solution for complex reasoning in search-augmented settings.

• We propose a two-stage fine-tuning framework that first applies SFT on a mixture of retrieval, reasoning, and decomposition datasets, followed by reinforcement fine-tuning using rewards solely from the final answer to train the decomposer and solver without intermediate supervision. This approach can be readily applied to LLMs with varying sizes (1.5B -32B as shown in our study) to enhance the multi-step reasoning ability of search-augmented LLMs.

• We conduct extensive evaluations of AceSearcher covering three tasks across ten public datasets.

Compared to strong baselines, including recent reasoning models and RL-enhanced search LLMs, AceSearcher demonstrates strong empirical performance with 7.6% gain on average. Moreover, AceSearcher demonstrates high parameter efficiency: the 1.5B variant matches the performance of models 10× larger on QA tasks, highlighting its suitability for low-resource settings.

2 Related Works Reasoning-intensive Search/Retrieval. Standard RAG pipelines often

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers4

Ask how each one uses it

Builds on39

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines