MSP: Probabilistically Consistent Multi-Scale Action Generation
Zhixuan Lin, Gengqi Liu, Chao Zheng, Gao Lin, Jindong Yu, Song Gao, Fei Wang
Abstract
In robotic imitation learning, accurately modeling the multimodality and temporal correlations of long-horizon action sequences remains challenging. Long-horizon tasks require preserving global task intent while executing precise low-level control; otherwise, local errors can accumulate and lead to failure. While recent coarse-to-fine autoregressive models have improved action generation, they struggle to maintain consistency across hierarchies, leading to suboptimal performance in long-horizon tasks. To address these shortcomings, we propose Probabilistically Consistent Multi-Scale Action Generation (MSP), a novel coarse-to-fine approach that promotes cross-scale consistency. MSP adopts a streamlined multi-scale design by directly downsampling in a continuous latent space. A scale-wise autoregressive Transformer is used to generate semantic conditions at each scale, which guide a lightweight MeanFlow model to capture multi-scale latent distributions, enabling probabilistically consistent refinement across scales. Through extensive simulation and real-world experiments, including long-horizon, multi-task, and few-shot generalization settings, we show that MSP outperforms existing coarse-to-fine methods, achieving state-of-the-art performance with high efficiency. Our code will be publicly available upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Mean Flows for One-step Generative ModelingZhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter et al.NeurIPS 2025 · 628 citations
- Finite Scalar Quantization: VQ-VAE Made SimpleFabian Mentzer, David Minnen, Eirikur Agustsson, Michael TschannenICLR 2024 · 442 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
Related papers
- Primary-Fine Decoupling for Action Generation in Robotic ImitationXiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu et al.ICLR 2026
- CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive PredictionZhefei Gong, Pengxiang Ding, Shangke Lyu, Siteng Huang et al.ICCV 2025 · 3 citations
- Next-Scale Autoregressive Models for Text-to-Motion GenerationZhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin ZhaoCVPR 2026 · 6 citations
- Masked Generative Policy for Robotic ControlLipeng Zhuang, Shiyu Fan, Florent P. Audonnet, Yingdong Ru et al.ICLR 2026 · 1 citation
- MAGE: Multi-scale Autoregressive Generation for Offline Reinforcement LearningChenxing Lin, Xinhui Gao, Haipeng Zhang, Xinran Li et al.ICLR 2026
