Improve Agents without Retraining: Parallel Tree Search with Off-Policy Correction
Gal Dalal, Assaf Hallak, Steven Dalton, Iuri Frosio, Shie Mannor, Gal Chechik
摘要
Tree Search (TS) is crucial to some of the most influential successes in reinforcement learning. Here, we tackle two major challenges with TS that limit its usability: distribution shift and scalability. We first discover and analyze a counter-intuitive phenomenon: action selection through TS and a pre-trained value function often leads to lower performance compared to the original pre-trained agent, even when having access to the exact state and reward in future steps. We show this is due to a distribution shift to areas where value estimates are highly inaccurate and analyze this effect using Extreme Value theory. To overcome this problem, we introduce a novel off-policy correction term that accounts for the mismatch between the pre-trained value and its corresponding TS policy by penalizing under-sampled trajectories. We prove that our correction eliminates the above mismatch and bound the probability of sub-optimal action selection. Our correction significantly improves pre-trained Rainbow agents without any further training, often more than doubling their scores on Atari games. Next, we address the scalability issue given by the computational complexity of exhaustive TS that scales exponentially with the tree depth. We introduce Batch-BFS: a GPU breadth-first search that advances all nodes in each depth of the tree simultaneously. Batch-BFS reduces runtime by two orders of magnitude and, beyond inference, enables also training with TS of depths that were not feasible before. We train DQN agents from scratch using TS and show improvement in several Atari games compared to both the original DQN and the more advanced Rainbow. The code for BCTS can be found in https://github.com/NVlabs/bcts . * Equal contribution (random order) 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Planning and Learning with Adaptive LookaheadAviv Rosenberg, Assaf Hallak, Shie Mannor, Gal Chechik 等AAAI 2023 · 被引用 12 次
- Generalised Policy Improvement with Geometric Policy CompositionShantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney 等ICML 2022 · 被引用 11 次
- SPO: Sequential Monte Carlo Policy OptimisationMatthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth 等NeurIPS 2024 · 被引用 8 次
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 被引用 7 次
- The Effective Horizon Explains Deep RL Performance in Stochastic EnvironmentsCassidy Laidlaw, Banghua Zhu, Stuart Russell, Anca D. DraganICLR 2024 · 被引用 5 次
它引用的顶会 Paper5
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain 等NeurIPS 2021 · 被引用 149 次
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau 等ICML 2021 · 被引用 137 次
- Accelerating Reinforcement Learning through GPU Atari EmulationSteven Dalton, Iuri FrosioNeurIPS 2020 · 被引用 50 次
- Learning to Simulate Dynamic Environments With GameGANSeung Wook Kim, Yuhao Zhou, Jonah Philion, Antonio Torralba 等CVPR 2020
相关 Paper
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- Online Pre-Training for Offline-to-Online Reinforcement LearningYongjae Shin, Jeonghye Kim, Whiyoung Jung, Sunghoon Hong 等ICML 2025
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang 等AAAI 2025 · 被引用 2 次
- Twice Sequential Monte Carlo for Tree SearchYaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan 等ICML 2026 · 被引用 2 次
- Online Tuning for Offline Decentralized Multi-Agent Reinforcement LearningJiechuan Jiang, Zongqing LuAAAI 2023 · 被引用 3 次
