DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
Changyeon Kim, Haeone Lee, Younggyo Seo, Kimin Lee, Yuke Zhu
Abstract
Offline reinforcement learning (RL) presents an attractive paradigm for training intelligent agents without expensive online interactions. However, current approaches still struggle with complex, long-horizon sequential decision making. In this work, we introduce DEtached value learning with Action Sequence (DEAS), a simple yet effective offline RL framework that leverages action sequences for value learning. These temporally extended actions provide richer information than single-step actions and can be interpreted through the options framework via semi-Markov decision process Q-learning, enabling reduction of the effective planning horizon by considering longer sequences at once. However, directly adopting such sequences in actor-critic algorithms introduces excessive value overestimation, which we address through detached value learning that steers value estimates toward in-distribution actions that achieve high return in the offline dataset. We demonstrate that DEAS consistently outperforms baselines on complex, long-horizon tasks from OGBench and can be applied to enhance the performance of large-scale Vision-Language-Action models that predict action sequences, significantly boosting performance in both RoboCasa Kitchen simulation tasks and real-world manipulation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee0f700e-fde4-4fc0-aac8-cce3126d69daCited by top-tier papers2
- Chunk-Guided Q-LearningGwanwoo Song, Kwanyoung Park, Youngwoon LeeICML 2026 · 2 citations
- QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RLXing Lei, Jincheng Wang, Xuetao Zhang, Donglin WangICML 2026
Builds on24
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
Related papers
- Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement LearningTian-Shuo Liu, Xu-Hui Liu, Ruifeng Chen, Lixuan Jin et al.ICLR 2025
- ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied TasksPhilip Schroeder, Ondrej Biza, Thomas Weng, Hongyin Luo et al.NeurIPS 2025 · 8 citations
- Scalable Offline Model-Based RL with Action ChunksKwanyoung Park, Seohong Park, Youngwoon Lee, Sergey LevineICLR 2026 · 12 citations
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li et al.ICML 2022 · 17 citations
- Option-aware Temporally Abstracted Value for Offline Goal-Conditioned Reinforcement LearningHongjoon Ahn, Heewoong Choi, Jisu Han, Taesup MoonNeurIPS 2025 · 22 citations
