Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns
Dong Tian, Onur Celik, Gerhard Neumann
Abstract
We introduce a sequence-conditioned critic for Soft Actor--Critic (SAC) that models trajectory context with a lightweight Transformer and trains on aggregated -step targets. Unlike prior approaches that (i) score state--action pairs in isolation or (ii) rely on actor-side action chunking to handle long horizons, our method strengthens the critic itself by conditioning on short trajectory segments and integrating multi-step returns without the need of importance sampling (IS). The resulting sequence-aware value estimates capture the critical temporal structure for extended-horizon and sparse-reward problems. On multiple benchmarks, we further show that freezing critic parameters for several steps makes our update compatible with CrossQ's core idea, enabling stable training without a target network. Despite its simplicity, a 2-layer Transformer with -- hidden units and a maximum update-to-data ratio (UTD) of , the approach consistently outperforms standard SAC and strong off-policy baselines, with particularly large gains on long-trajectory control. These results highlight the value of sequence modeling and -step bootstrapping on the critic side for long-horizon reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8be517ec-024a-4abc-9629-160f5aab8593Cited by top-tier papers3
- Reinforcement Learning with Action ChunkingQiyang Li, Zhiyuan Zhou, Sergey LevineNeurIPS 2025 · 114 citations
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 19 citations
- DEAS: DEtached value learning with Action Sequence for Scalable Offline RLChangyeon Kim, Haeone Lee, Younggyo Seo, Kimin Lee et al.ICLR 2026 · 9 citations
Builds on17
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
Related papers
- Decision S4: Efficient Sequence-Based RL via State Spaces LayersShmuel Bar-David, Itamar Zimerman, Eliya Nachmani, Lior WolfICLR 2023 · 3 citations
- TOP-ERL: Transformer-based Off-Policy Episodic Reinforcement LearningGe Li, Dong Tian, Hongyi Zhou, Xinkai Jiang et al.ICLR 2025
- Multi-timescale Reinforcement Learning by Value ReconstructionZhan Su, Peixi Peng, Xinyu Hu, Cong Li et al.ICML 2026
- Q-value Regularized Transformer for Offline Reinforcement LearningShengchao Hu, Ziqing Fan, Chaoqin Huang, Li Shen et al.ICML 2024 · 34 citations
- CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and SimplicityAditya Bhatt, Daniel Palenicek, Boris Belousov, Max Argus et al.ICLR 2024 · 106 citations
