Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers
Xin Chen, Feng Jiang, Yiqian Zhang, Hardy Chen, Shuo Yan, Wenya Xie, Min Yang, Shujian Huang
Abstract
Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain fundamentally limited by a blind self-thinking paradigm: performing extensive internal reasoning even when critical information is missing or ambiguous. We propose Proactive Interactive Reasoning (PIR), a new reasoning paradigm that transforms LLMs from passive solvers into proactive inquirers that interleave reasoning with clarification. Unlike existing search- or tool-based frameworks that primarily address knowledge uncertainty by querying external environments, PIR targets premise- and intent-level uncertainty through direct interaction with the user. PIR is implemented via two core components: (1) an uncertainty-aware supervised fine-tuning procedure that equips models with interactive reasoning capability, and (2) a user-simulator-based policy optimization framework driven by a composite reward that aligns model behavior with user intent. Extensive experiments on mathematical reasoning, code generation, and document editing demonstrate that PIR consistently outperforms strong baselines, achieving up to 32.70% higher accuracy, 22.90% higher pass rate, and 41.36 BLEU improvement, while reducing nearly half of the reasoning computation and unnecessary interaction turns. Further reliability evaluations on factual knowledge, question answering, and missing-premise scenarios confirm the strong generalization and robustness of PIR. Model and code are publicly available at: https://github.com/SUAT-AIRI/Proactive-Interactive-R1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a5d5165-d1e0-42e8-8af6-16aaa12b8cc3Builds on9
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement LearningHaoran Luo, Haihong E, Guanting Chen, Qika Lin et al.ICML 2026 · 50 citations
- Generating Multi-turn Clarification for Web Information SeekingZiliang Zhao, Zhicheng DouWWW 2024 · 15 citations
- Interactive Text GenerationFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng et al.EMNLP 2023 · 7 citations
Related papers
- Clarifying Ambiguities: on the Role of Ambiguity Types in Prompting Methods for Clarification GenerationAnfu Tang, Laure Soulier, Vincent GuigueSIGIR 2025 · 6 citations
- Rethinking Chain-of-Thought from the Perspective of Self-TrainingZongqian Wu, Baoduo Xu, Ruochen Cui, Mengmeng Zhan et al.ICML 2025
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu et al.EMNLP 2023 · 184 citations
- SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep ReasoningShuzheng Gao, Chaozheng Wang, Cuiyun Gao, Michael R. LyuICSE 2026 · 1 citation
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 13 citations
