Efficient Planning in a Compact Latent Action Space
Zhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li, Tim Rocktäschel, Edward Grefenstette, Yuandong Tian
Abstract
Planning-based reinforcement learning has shown strong performance in tasks in discrete and low-dimensional continuous action spaces. However, planning usually brings significant computational overhead for decision-making, and scaling such methods to high-dimensional action spaces remains challenging. To advance efficient planning for high-dimensional continuous control, we propose Trajectory Autoencoding Planner (TAP), which learns low-dimensional latent action codes with a state-conditional VQ-VAE. The decoder of the VQ-VAE thus serves as a novel dynamics model that takes latent actions and current state as input and reconstructs long-horizon trajectories. During inference time, given a starting state, TAP searches over discrete latent actions to find trajectories that have both high probability under the training distribution and high predicted cumulative reward. Empirical evaluation in the offline RL setting demonstrates low decision latency which is indifferent to the growing raw action dimensionality. For Adroit robotic hand manipulation tasks with high-dimensional continuous action space, TAP surpasses existing model-based methods by a large margin and also beats strong model-free actor-critic baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers41
- HIQL: Offline Goal-Conditioned RL with Latent States as ActionsSeohong Park, Dibya Ghosh, Benjamin Eysenbach, Sergey LevineNeurIPS 2023 · 173 citations
- QueST: Self-Supervised Skill Abstractions for Learning Continuous ControlAtharva Mete, Haotian Xue, Albert Wilcox, Yongxin Chen et al.NeurIPS 2024 · 76 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- The Wisdom of Hindsight Makes Language Models Better Instruction FollowersTianjun Zhang, Fangchen Liu, Justin Wong, Pieter Abbeel et al.ICML 2023 · 65 citations
- Masked Trajectory Models for Prediction, Representation, and ControlPhilipp Wu, Arjun Majumdar, Kevin Stone, Yixin Lin et al.ICML 2023 · 57 citations
Builds on17
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- Scalable Decision-Making in Stochastic Environments through Learned Temporal AbstractionBaiting Luo, Ava Pettet, Aron Laszka, Abhishek Dubey et al.ICLR 2025
- HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action RepresentationBoyan Li, Hongyao Tang, Yan Zheng, Jianye Hao et al.ICLR 2022 · 79 citations
- TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic ManipulationSizhe Zhao, Shengping Zhang, Shuo Yang, Weiyu Zhao et al.ICML 2026 · 1 citation
- Learning High-Frequency Continuous Action Chunks in Latent SpaceKunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao et al.ICML 2026
- Character controllers using motion VAEsHung Yu Ling, Fabio Zinno, George Cheng, Michiel van de PanneSIGGRAPH 2020 · 261 citations
