Planning and Learning with Adaptive Lookahead
Aviv Rosenberg, Assaf Hallak, Shie Mannor, Gal Chechik, Gal Dalal
Abstract
Some of the most powerful reinforcement learning frameworks use planning for action selection. Interestingly, their planning horizon is either fixed or determined arbitrarily by the state visitation history. Here, we expand beyond the naive fixed horizon and propose a theoretically justified strategy for adaptive selection of the planning horizon as a function of the state-dependent value estimate. We propose two variants for lookahead selection and analyze the trade-off between iteration count and computational complexity per iteration. We then devise a corresponding deep Q-network algorithm with an adaptive tree search horizon. We separate the value estimation per depth to compensate for the off-policy discrepancy between depths. Lastly, we demonstrate the efficacy of our adaptive lookahead method in a maze environment and Atari. * Research conducted while the author was an intern at Nvidia Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 567ca32a-4022-410d-accf-96cc5d95fddeCited by top-tier papers1
Ask how each one uses itBuilds on3
- Online Planning with Lookahead PoliciesYonathan Efroni, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2020 · 25 citations
- On Covariate Shift of Latent Confounders in Imitation and Reinforcement LearningGuy Tennenholtz, Assaf Hallak, Gal Dalal, Shie Mannor et al.ICLR 2022 · 16 citations
- Improve Agents without Retraining: Parallel Tree Search with Off-Policy CorrectionGal Dalal, Assaf Hallak, Steven Dalton, Iuri Frosio et al.NeurIPS 2021 · 11 citations
Related papers
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 25 citations
- Stop! Planner Time: Metareasoning for Probabilistic Planning Using Learned Performance ProfilesMatthew Budd, Bruno Lacerda, Nick HawesAAAI 2024 · 2 citations
- Learning the Target Network in Function SpaceKavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin et al.ICML 2024 · 3 citations
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 22 citations
- The Effective Horizon Explains Deep RL Performance in Stochastic EnvironmentsCassidy Laidlaw, Banghua Zhu, Stuart Russell, Anca D. DraganICLR 2024 · 5 citations
