CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction
Zhefei Gong, Pengxiang Ding, Shangke Lyu, Siteng Huang, Mingyang Sun, Wei Zhao, Zhaoxin Fan, Donglin Wang
摘要
In robotic visuomotor policy learning, diffusion-based models have achieved significant success in improving the accuracy of action trajectory generation compared to traditional autoregressive models. However, they suffer from inefficiency due to multiple denoising steps and limited flexibility from complex constraints. In this paper, we introduce Coarse-to-Fine AutoRegressive Policy (CARP), a novel paradigm for visuomotor policy learning that redefines the autoregressive action generation process as a coarse-to-fine, next-scale approach. CARP decouples action generation into two stages: first, an action autoencoder learns multi-scale representations of the entire action sequence; then, a GPT-style transformer refines the sequence prediction through a coarse-to-fine autoregressive process. This straightforward and intuitive approach produces highly accurate and smooth actions, matching or even surpassing the performance of diffusion-based policies while maintaining efficiency on par with autoregressive policies. We conduct extensive evaluations across diverse settings, including single-task and multi-task scenarios on state-based and image-based simulation benchmarks, as well as real-world tasks. CARP achieves competitive success rates, with up to a 10% improvement, and delivers 10x faster inference compared to state-of-the-art policies, establishing a high-performance, efficient, and flexible paradigm for action generation in robotic tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- VITA: Vision-to-Action Flow Matching PolicyDechen Gao, BOQI ZHAO, Andrew Lee, Ian Chuang 等ICLR 2026 · 被引用 27 次
- Dense Policy: Bidirectional Autoregressive Learning of ActionsYue Su, Xinyu Zhan, Hongjie Fang, Han Xue 等ICCV 2025 · 被引用 23 次
- FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous TokensYiming Zhong, Yumeng Liu, Chuyang Xiao, Zemin Yang 等NeurIPS 2025 · 被引用 16 次
- HDP: Triply‑Hierarchical Diffusion Policy for Visuomotor LearningYiyang Lu, Yufeng Tian, Zhecheng Yuan, Xianbang Wang 等ICLR 2026 · 被引用 10 次
- SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World ModelsSen Wang, Jingyi Tian, Le Wang, Zhimin Liao 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Masked Generative Policy for Robotic ControlLipeng Zhuang, Shiyu Fan, Florent P. Audonnet, Yingdong Ru 等ICLR 2026 · 被引用 1 次
- Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationWenbo Zhang, Tianrun Hu, Hanbo Zhang, Yanyuan Qiao 等NeurIPS 2025 · 被引用 19 次
- Prediction with Action: Visual Policy Learning via Joint Denoising ProcessYanjiang Guo, Yucheng Hu, Jianke Zhang, Yen-Jen Wang 等NeurIPS 2024 · 被引用 93 次
- MSP: Probabilistically Consistent Multi-Scale Action GenerationZhixuan Lin, Gengqi Liu, Chao Zheng, Gao Lin 等ICML 2026
- Quantization-Free Autoregressive Action TransformerZiyad Sheebaelhamd, Michael Tschannen, Michael Muehlebach, Claire VernadeNeurIPS 2025 · 被引用 4 次
