BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces
Matthew Landers, Taylor W. Killian, Hugo Barnes, Tom Hartvigsen, Afsaneh Doryab
Abstract
Offline reinforcement learning in high-dimensional, discrete action spaces is challenging due to the exponential scaling of the joint action space with the number of sub-actions and the complexity of modeling sub-action dependencies. Existing methods either exhaustively evaluate the action space, making them computationally infeasible, or factorize Q-values, failing to represent joint sub-action effects. We propose Branch Value Estimation (BraVE), a value-based method that uses tree-structured action traversal to evaluate a linear number of joint actions while preserving dependency structure. BraVE outperforms prior offline RL methods by up to in environments with over four million actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a00f397e-816c-48ff-9040-0487154cdf88Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Reinforcement Learning with Combinatorial Actions: An Application to Vehicle RoutingArthur Delarue, Ross Anderson, Christian TjandraatmadjaNeurIPS 2020 · 127 citations
Related papers
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez et al.NeurIPS 2022 · 63 citations
- Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement LearningJianzhun Shao, Yun Qu, Chen Chen, Hongchang Zhang et al.NeurIPS 2023 · 56 citations
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 8 citations
- Inferring DQN structure for high-dimensional continuous controlAndrey Sakryukin, Chedy Raïssi, Mohan S. KankanhalliICML 2020 · 8 citations
- Offline RL Without Off-Policy EvaluationDavid Brandfonbrener, Will Whitney, Rajesh Ranganath, Joan BrunaNeurIPS 2021 · 217 citations
