Test-Time Deep Thinking to Explore Implicit Rules
Wentong Chen, Xin Cong, Zhong Zhang, Yaxi Lu, Siyuan Zhao, Yesai Wu, Qinyu Luo, Haotian Chen, Yankai Lin, Zhiyuan Liu, Maosong Sun
Abstract
With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in environments governed by implicit rules—hidden constraints that cannot be observed directly and must be inferred through interaction. This causes agents to fall into repetitive trial-and-error loops, ultimately leading to task failure. To address this challenge, we propose Test-Time Exploration (TTExplore), a framework where a thinker component analyzes interaction history to infer these implicit rules and guide an actor. Effective exploration in this setting critically depends on the reasoning ability of the thinker. However, evaluating deep reasoning trajectories is inherently unstable and difficult, which poses a major obstacle to effective training. To overcome this issue, we introduce a novel and stable reinforcement learning pipeline. The core idea is to use accurate task-level scores as indirect rewards to bypass the difficulty of evaluating intermediate reasoning, and to retain only a single thinking node per trajectory to alleviate reward sparsity. Using this pipeline, we train a specialized 7B model, Exp-Thinker. Experiments on five text-based embodied tasks show that TTExplore equipped with Exp-Thinker improves baseline agent performance by an average of 14-19 points, demonstrating the effectiveness of explicitly reasoning about implicit rules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98ca10d7-7665-49d8-b49c-be894960ec6dBuilds on19
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin et al.AAAI 2024 · 484 citations
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksBill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman et al.NeurIPS 2023 · 244 citations
Related papers
- Learning Structured Reasoning via Tractable Trajectory ControlPo-Nien Kung, Zhen Yang, Jeffrey Luo, Cheng-Fu Yang et al.ICML 2026
- Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningShu Wu, Chenxing Li, Wenfu Wang, Hao Zhang et al.AAAI 2026 · 4 citations
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsAmrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer et al.ICLR 2026 · 66 citations
- Rectifying LLM Thought from Lens of OptimizationJunnan Liu, Hongwei Liu, Songyang Zhang, Kai ChenICLR 2026 · 3 citations
- T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference ScalingZhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang et al.ICML 2025
