IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
Haohao Luo, Zexi Li, Yuexiang Xie, Wenhao Zhang, Yaliang Li, Ying Shen
Abstract
Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computationally expensive and time-consuming, creating an autonomy-interaction dilemma: high autonomy on ambiguous user queries often leads to prolonged execution with unsatisfactory outcomes. To address this, we propose IntentRL, a framework that trains proactive agents to clarify latent user intents before starting long-horizon research. To overcome the scarcity of open-ended research data, we introduce a scalable pipeline that expands a few seed samples into high-quality dialogue turns via a shallow-to-deep intent refinement graph. We further adopt a two-stage reinforcement learning (RL) strategy: Stage I applies RL on offline dialogues to efficiently learn general user-interaction behavior, while Stage II uses the trained agent and a user simulator for online rollouts to strengthen adaptation to diverse user feedback. Extensive experiments show that IntentRL significantly improves both intent hit rate and downstream task performance, outperforming the built-in clarify modules of closed-source DR agents and proactive LLM baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db752aa7-57c0-45ae-962b-6976a7388b3cCited by top-tier papers1
Ask how each one uses itBuilds on10
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- Deep Research Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded TasksHaiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu et al.AAAI 2026 · 20 citations
- Towards Personalized Deep Research: Benchmarks and EvaluationsYuan Liang, Jiaxian Li, Yuqing Wang, Piaohong Wang et al.ICLR 2026 · 13 citations
- Grounded in Reality: Learning and Deploying Proactive LLM from Offline LogsFei Wei, Daoyuan Chen, Ce Wang, Yilun Huang et al.ICML 2026 · 2 citations
- Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-TrainingMaximillian Chen, Ruoxi Sun, Tomas Pfister, Sercan Ö. ArikICLR 2025 · 1 citation
Related papers
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language ModelsWei Liu, Peijie Yu, Michele Orini, Yali Du et al.ICML 2026 · 2 citations
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world EnvironmentsYuxiang Zheng, Dayuan Fu, Xiangkun Hu, Xiaojie Cai et al.EMNLP 2025 · 8 citations
- Learning to Retrieve from Agent TrajectoriesYuqi Zhou, Sunhao Dai, Changle Qu, Liang Pang et al.SIGIR 2026
- FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based AgentsChiwei Zhu, Benfeng Xu, Mingxuan Du, Shaohan Wang et al.ACL 2026 · 2 citations
- Reinforcement Learning with Evolving Rubrics for Deep ResearchRulin Shao, Akari Asai, Shannon Shen, Hamish Ivison et al.ICML 2026 · 78 citations
