STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction
Jinhao Li, Yuxuan Cong, Yingqiao Wang, Hao Xia, Shan Huang, Yijia Zhang, Ningyi Xu, Guohao Dai
Abstract
Diffusion policies have recently emerged as a powerful paradigm for visuomotor control in robotic manipulation due to their ability to model the distribution of action sequences and capture multimodality. However, iterative denoising leads to substantial inference latency, limiting control frequency in real-time closed-loop systems. Existing acceleration methods either reduce sampling steps, bypass diffusion through direct prediction, or reuse past actions, but often struggle to jointly preserve action quality and achieve consistently low latency. In this work, we propose STEP, a lightweight spatiotemporal consistency prediction mechanism to construct high-quality warmstart actions that are both distributionally close to the target action and temporally consistent, without compromising the generative capability of the original diffusion policy. Then, we propose a velocity-aware perturbation injection mechanism that adaptively modulates actuation excitation based on temporal action variation to prevent execution stall especially for real-world tasks. We further provide a theoretical analysis showing that the proposed prediction induces a locally contractive mapping, ensuring convergence of action errors during diffusion refinement. We conduct extensive evaluations on nine simulated benchmarks and two real-world tasks. Notably, STEP with 2 steps can achieve an average 21.6% and 27.5% higher success rate than BRIDGER and DDIM on the RoboMimic benchmark and real-world tasks, respectively. These results demonstrate that STEP consistently advances the Pareto frontier of inference latency and success rate over existing methods. The code is publicly available at https://github.com/Kimho666/STEP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen et al.NeurIPS 2022 · 2,653 citations
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation SkillsJiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling et al.ICLR 2023 · 21 citations
- MP1: MeanFlow Tames Policy Learning in 1-step for Robotic ManipulationJuyi Sheng, Ziyi Wang, Peiming Li, Mengyuan LiuAAAI 2026 · 18 citations
Related papers
- On-Device Diffusion Transformer Policy for Efficient Robot ManipulationYiming Wu, Huan Wang, Zhenghao Chen, Jianxin Pang et al.ICCV 2025 · 1 citation
- Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential EquationZhaoyang Liu, Mokai Pan, Zhongyi Wang, Kaizhen Zhu et al.ICML 2026 · 2 citations
- One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion DistillationZhendong Wang, Max Li, Ajay Mandlekar, Zhenjia Xu et al.ICML 2025
- Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsHao Ren, Yiming Zeng, Zetong Bi, Zhaoliang Wan et al.CVPR 2025
- Falcon: Fast Visuomotor Policies via Partial DenoisingHaojun Chen, Minghao Liu, Chengdong Ma, Xiaojian Ma et al.ICML 2025
