Sub-Goal Trees a Framework for Goal-Based Reinforcement Learning
Tom Jurgenson, Or Avner, Edward Groshev, Aviv Tamar
摘要
Many AI problems, in robotics and other domains, are goal-based, essentially seeking trajectories leading to various goal states. Reinforcement learning (RL), building on Bellman's optimality equation, naturally optimizes for a single goal, yet can be made multi-goal by augmenting the state with the goal. Instead, we propose a new RL framework, derived from a dynamic programming equation for the all pairs shortest path (APSP) problem, which naturally solves multi-goal queries. We show that this approach has computational benefits for both standard and approximate dynamic programming. Interestingly, our formulation prescribes a novel protocol for computing a trajectory: instead of predicting the next state given its predecessor, as in standard RL, a goal-conditioned trajectory is constructed by first predicting an intermediate state between start and goal, partitioning the trajectory into two. Then, recursively, predicting intermediate points on each sub-segment, until a complete trajectory is obtained. We call this trajectory structure a sub-goal tree. Building on it, we additionally extend the policy gradient methodology to recursively predict sub-goals, resulting in novel goal-based algorithms. Finally, we apply our method to neural motion planning, where we demonstrate significant improvements compared to standard RL on navigating a 7-DoF robot arm between obstacles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 被引用 183 次
- BAKU: An Efficient Transformer for Multi-Task Policy LearningSiddhant Haldar, Zhuoran Peng, Lerrel PintoNeurIPS 2024 · 被引用 120 次
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou 等NeurIPS 2020 · 被引用 96 次
- World Model as a Graph: Learning Latent Landmarks for PlanningLunjun Zhang, Ge Yang, Bradly C. StadieICML 2021 · 被引用 90 次
- CO-PILOT: COllaborative Planning and reInforcement Learning On sub-Task curriculumShuang Ao, Tianyi Zhou, Guodong Long, Qinghua Lu 等NeurIPS 2021 · 被引用 23 次
相关 Paper
- BT-Tree: A Reinforcement Learning Based Index for Big Trajectory DataTu Gu, Kaiyu Feng, Jingyi Yang, Gao Cong 等SIGMOD 2025 · 被引用 3 次
- Hierarchical Imitation Learning with Vector Quantized ModelsKalle Kujanpää, Joni Pajarinen, Alexander IlinICML 2023 · 被引用 17 次
- Chain-of-Goals Hierarchical Policy for Long-Horizon Offline Goal-Conditioned RLJinwoo Choi, Sang-Hyun Lee, Seung-Woo SeoICML 2026 · 被引用 3 次
- Efficient and Effective Similar Subtrajectory Search with Deep Reinforcement LearningZheng Wang, Cheng Long, Gao Cong, Yiding LiuVLDB 2020 · 被引用 29 次
- Neuro-algorithmic Policies Enable Fast Combinatorial GeneralizationMarin Vlastelica P., Michal Rolínek, Georg MartiusICML 2021 · 被引用 17 次
