Planning and Learning with Adaptive Lookahead
Aviv Rosenberg, Assaf Hallak, Shie Mannor, Gal Chechik, Gal Dalal
摘要
Some of the most powerful reinforcement learning frameworks use planning for action selection. Interestingly, their planning horizon is either fixed or determined arbitrarily by the state visitation history. Here, we expand beyond the naive fixed horizon and propose a theoretically justified strategy for adaptive selection of the planning horizon as a function of the state-dependent value estimate. We propose two variants for lookahead selection and analyze the trade-off between iteration count and computational complexity per iteration. We then devise a corresponding deep Q-network algorithm with an adaptive tree search horizon. We separate the value estimation per depth to compensate for the off-policy discrepancy between depths. Lastly, we demonstrate the efficacy of our adaptive lookahead method in a maze environment and Atari. * Research conducted while the author was an intern at Nvidia Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Online Planning with Lookahead PoliciesYonathan Efroni, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2020 · 被引用 25 次
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor 等ICLR 2022 · 被引用 16 次
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy CorrectionGal Dalal, Assaf Hallak, Steven Dalton, Iuri Frosio 等NeurIPS 2021 · 被引用 11 次
相关 Paper
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 被引用 25 次
- Stop! Planner Time: Metareasoning for Probabilistic Planning Using Learned Performance ProfilesMatthew Budd, Bruno Lacerda, Nick HawesAAAI 2024 · 被引用 2 次
- Learning the Target Network in Function SpaceKavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin 等ICML 2024 · 被引用 3 次
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- The Effective Horizon Explains Deep RL Performance in Stochastic EnvironmentsCassidy Laidlaw, Banghua Zhu, Stuart Russell, Anca D. DraganICLR 2024 · 被引用 5 次
