Dense Policy: Bidirectional Autoregressive Learning of Actions
Yue Su, Xinyu Zhan, Hongjie Fang, Han Xue, Hao-Shu Fang, Yong-Lu Li, Cewu Lu, Lixin Yang
Abstract
Mainstream visuomotor policies predominantly rely on generative models for holistic action prediction, while current autoregressive policies, predicting the next token or chunk, have shown suboptimal results. This motivates a search for more effective learning methods to unleash the potential of autoregressive policies for robotic manipulation. This paper introduces a bidirectionally expanded learning approach, termed Dense Policy, to establish a new paradigm for autoregressive policies in action prediction. It employs a lightweight encoder-only architecture to iteratively unfold the action sequence from an initial single frame into the target sequence in a coarse-to-fine manner with logarithmic-time inference. Extensive experiments validate that our dense policy has superior autoregressive learning capabilities and can surpass existing holistic generative policies. Our model, data, and code are available at: https://selen-suyue.github.io/DspNet/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- VITA: Vision-to-Action Flow Matching PolicyDechen Gao, BOQI ZHAO, Andrew Lee, Ian Chuang et al.ICLR 2026 · 27 citations
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation PriorsZhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu et al.ICLR 2026 · 27 citations
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
- FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous TokensYiming Zhong, Yumeng Liu, Chuyang Xiao, Zemin Yang et al.NeurIPS 2025 · 16 citations
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion RepresentationsShichao Fan, Kun Wu, Zhengping Che, Xinhua Wang et al.ICML 2026 · 16 citations
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
Related papers
- CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive PredictionZhefei Gong, Pengxiang Ding, Shangke Lyu, Siteng Huang et al.ICCV 2025 · 3 citations
- Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationWenbo Zhang, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao et al.NeurIPS 2025 · 19 citations
- Learning Foresightful Dense Visual Affordance for Deformable Object ManipulationRuihai Wu, Chuanruo Ning, Hao DongICCV 2023 · 45 citations
- Masked Generative Policy for Robotic ControlLipeng Zhuang, Shiyu Fan, Florent P. Audonnet, Yingdong Ru et al.ICLR 2026 · 1 citation
- FASTer: Toward Powerful and Efficient Autoregressive Vision-Language-Action Models with Learnable Action Tokenizer and Block-wise DecodingYicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye et al.ICLR 2026
