Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
Younggyo Seo, Pieter Abbeel
Abstract
Predicting a sequence of actions has been crucial in the success of recent behavior cloning algorithms in robotics. Can similar ideas improve reinforcement learning (RL)? We answer affirmatively by observing that incorporating action sequences when predicting ground-truth return-to-go leads to lower validation loss. Motivated by this, we introduce Coarse-to-fine Q-Network with Action Sequence (CQN-AS), a novel value-based RL algorithm that learns a critic network that outputs Q-values over a sequence of actions, i.e., explicitly training the value function to learn the consequence of executing action sequences. Our experiments show that CQN-AS outperforms several baselines on a variety of sparse-reward humanoid control and tabletop manipulation tasks from BiGym and RLBench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdfe3ea4-d026-4a78-9e2a-2ee10b697332Cited by top-tier papers3
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 114 citations
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 19 citations
- Offline Reinforcement Learning with Universal Horizon ModelsHojun Chung, Junseo Lee, Songhwai OhICML 2026 · 1 citation
Builds on9
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-TuningMitsuhiko Nakamoto, Simon Zhai, Anikait Singh, Max Sobol Mark et al.NeurIPS 2023 · 296 citations
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz et al.ICML 2024 · 286 citations
Related papers
- Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-NetworkJijia Liu, Feng Gao, Qingmin Liao, Chao Yu et al.ICML 2025
- Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse RewardJiarui Yang, Bin Zhu, Jingjing Chen, Yu-Gang JiangAAAI 2026 · 3 citations
- DEAS: DEtached value learning with Action Sequence for Scalable Offline RLChangyeon Kim, Haeone Lee, Younggyo Seo, Kimin Lee et al.ICLR 2026 · 9 citations
- Imitation Learning for Human Pose PredictionBorui Wang, Ehsan Adeli, Hsu-Kuang Chiu, De-An Huang et al.ICCV 2019 · 110 citations
- Value-aligned Behavior Cloning for Offline Reinforcement Learning via Bi-level OptimizationXingyu Jiang, Ning Gao, Xiuhui Zhang, Hongkun Dou et al.ICLR 2025
