Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
Matthew Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh Doryab
Abstract
Reinforcement learning in combinatorial action spaces requires searching over exponentially many joint actions to simultaneously select multiple sub-actions that form coherent combinations. Existing approaches either simplify policy learning by assuming independence across sub-actions, which often yields incoherent or invalid actions when coordination is required, or attempt to learn action structure and control jointly, which is slow and unstable. We introduce Structured Policy Initialization (SPIN), a two-stage framework that first pre-trains an Action Structure Model (ASM) to capture the manifold of valid actions, then freezes this representation and trains lightweight policy heads for control. On challenging DM Control benchmarks, SPIN improves average return by up to over the state of the art while reducing time to convergence by up to .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb1c587c-511a-4302-8b6c-8be30ec9aa83Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial ActionsLingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti et al.ICML 2026 · 1 citation
- Explore to Learn: Latent Exploration Through Disentangled Synergy Patterns for Reinforcement Learning in Overactuated ControlYiming Wang, Kaiyan Zhao, Xu Li, Yan Li et al.AAAI 2026 · 1 citation
- Solving Challenging Dexterous Manipulation Tasks With Trajectory Optimisation and Reinforcement LearningHenry Charlesworth, Giovanni MontanaICML 2021 · 30 citations
- DynSyn: Dynamical Synergistic Representation for Efficient Learning and Control in Overactuated Embodied SystemsKaibo He, Chenhui Zuo, Chengtian Ma, Yanan SuiICML 2024 · 19 citations
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 112 citations
