Structure-Induced Information for Rerooting Levin Tree Search
Jake Tuero, Michael Buro, Laurent Orseau, Levi Lelis
摘要
Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but often relies on explicit subgoal generation that can incur substantial overhead and hinders scalability. In this paper, we overcome these limitations by using a learned ``rerooter'' through the recently-introduced algorithm. A rerooter implicitly decomposes the problem into soft subtasks. While previous work focused on the formal guarantees for given or handcrafted rerooters, in this work we propose three rerooter designs: (i) a clustering-based rerooter that exploits global state-space structure, (ii) a heuristic-based rerooter that leverages learned cost-to-go estimates, and (iii) a hybrid that combines both signals. Our framework avoids having to explicitly reconstruct and reason over generated subgoals, thereby enabling scalable allocation of search effort with significantly lower computational overhead. Empirically, our rerooting-based methods scale to complex environments where subgoal-based policy tree search fails, and achieve state-of-the-art online training efficiency on the domains tested.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Subgoal Search For Complex Reasoning TasksKonrad Czechowski, Tomasz Odrzygózdz, Marek Zbysinski, Michal Zawalski 等NeurIPS 2021 · 被引用 41 次
- Policy-Guided Heuristic Search with GuaranteesLaurent Orseau, Levi H. S. LelisAAAI 2021 · 被引用 30 次
- Hierarchical Imitation Learning with Vector Quantized ModelsKalle Kujanpää, Joni Pajarinen, Alexander IlinICML 2023 · 被引用 17 次
- Creating Multi-Level Skill Hierarchies in Reinforcement LearningJoshua B. Evans, Özgür SimsekNeurIPS 2023 · 被引用 15 次
- Hybrid Search for Efficient Planning with Completeness GuaranteesKalle Kujanpää, Joni Pajarinen, Alexander IlinNeurIPS 2023 · 被引用 7 次
相关 Paper
- Learning Rational Subgoals from Demonstrations and InstructionsZhezheng Luo, Jiayuan Mao, Jiajun Wu, Tomás Lozano-Pérez 等AAAI 2023 · 被引用 5 次
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
- Policy Gradient with Tree ExpansionGal Dalal, Assaf Hallak, Gugan Thoppe, Shie Mannor 等ICML 2025
- AlphaRouter: Token-level Routing Between SLM and LLM with Reinforcement Learning and Tree SearchSiteng Liao, Yuzhu Liang, Hengzhong Rao, Xizhao Luo 等ICML 2026
- Scalable Option Learning in High-Throughput EnvironmentsMikael Henaff, Scott Fujimoto, Michael Matthews, Michael RabbatICML 2026 · 被引用 5 次
