Subgoal Search For Complex Reasoning Tasks
Konrad Czechowski, Tomasz Odrzygózdz, Marek Zbysinski, Michal Zawalski, Krzysztof Olejnik, Yuhuai Wu, Lukasz Kucinski, Piotr Milos
摘要
Humans excel in solving complex reasoning tasks through a mental process of moving from one idea to a related one. Inspired by this, we propose Subgoal Search (kSubS) method. Its key component is a learned subgoal generator that produces a diversity of subgoals that are both achievable and closer to the solution. Using subgoals reduces the search space and induces a high-level search graph suitable for efficient planning. In this paper, we implement kSubS using a transformer-based subgoal module coupled with the classical best-first search framework. We show that a simple approach of generating -th step ahead subgoals is surprisingly efficient on three challenging domains: two popular puzzle games, Sokoban and the Rubik's Cube, and an inequality proving benchmark INT. kSubS achieves strong results including state-of-the-art on INT within a modest computational budget.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement LearningJunsu Kim, Younggyo Seo, Jinwoo ShinNeurIPS 2021 · 被引用 90 次
- Proving Theorems RecursivelyHaiming Wang, Huajian Xin, Zhengying Liu, Wenda Li 等NeurIPS 2024 · 被引用 34 次
- Formal Mathematics Statement Curriculum LearningStanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys 等ICLR 2023 · 被引用 24 次
- Hierarchical Imitation Learning with Vector Quantized ModelsKalle Kujanpää, Joni Pajarinen, Alexander IlinICML 2023 · 被引用 17 次
- Contrastive Representations for Temporal ReasoningAlicja Ziarko, Michal Bortkiewicz, Michal Zawalski, Benjamin Eysenbach 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper8
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 被引用 183 次
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 被引用 152 次
- Maximum Entropy Gain Exploration for Long Horizon Multi-goal Reinforcement LearningSilviu Pitis, Harris Chan, Stephen Zhao, Bradly C. Stadie 等ICML 2020 · 被引用 145 次
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical PredictorsKarl Pertsch, Oleh Rybkin, Frederik Ebert, Shenghao Zhou 等NeurIPS 2020 · 被引用 96 次
- World Model as a Graph: Learning Latent Landmarks for PlanningLunjun Zhang, Ge Yang, Bradly C. StadieICML 2021 · 被引用 90 次
相关 Paper
- Fast and Precise: Adjusting Planning Horizon with Adaptive Subgoal SearchMichal Zawalski, Michal Tyrolski, Konrad Czechowski, Tomasz Odrzygózdz 等ICLR 2023
- Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient DescentTong Yang, Yu Huang, Yingbin Liang, Yuejie ChiNeurIPS 2025 · 被引用 8 次
- INT: An Inequality Benchmark for Evaluating Generalization in Theorem ProvingYuhuai Wu, Albert Q. Jiang, Jimmy Ba, Roger Baker GrosseICLR 2021 · 被引用 60 次
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio 等ICLR 2024 · 被引用 21 次
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 被引用 69 次
