Real-Time Execution of Action Chunking Flow Policies
Kevin Black, Manuel Y. Galliker, Sergey Levine
Abstract
Modern AI systems, especially those interacting with the physical world, increasingly require real-time performance. However, the high latency of state-of-the-art generalist models, including recent vision-language action models (VLAs), poses a significant challenge. While action chunking has enabled temporal consistency in high-frequency control tasks, it does not fully address the latency problem, leading to pauses or out-of-distribution jerky movements at chunk boundaries. This paper presents a novel inference-time algorithm that enables smooth asynchronous execution of action chunking policies. Our method, real-time chunking (RTC), is applicable to any diffusion- or flow-based VLA out of the box with no re-training. It generates the next action chunk while executing the current one,"freezing"actions guaranteed to execute and"inpainting"the rest. To test RTC, we introduce a new benchmark of 12 highly dynamic tasks in the Kinetix simulator, as well as evaluate 6 challenging real-world bimanual manipulation tasks. Results demonstrate that RTC is fast, performant, and uniquely robust to inference delay, significantly improving task throughput and enabling high success rates in precise tasks such as lighting a match even in the presence of significant latency. See https://pi.website/research/real_time_chunking for videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1c8dc2a-7f07-47e2-a964-2d425c0d5082Cited by top-tier papers19
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 163 citations
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung et al.ICLR 2026 · 60 citations
- Real-Time Robot Execution with Masked Action ChunkingHaoxuan Wang, Gengyu Zhang, Yan Yan, Yuzhang Shang et al.ICLR 2026 · 24 citations
- Decoupled Q-ChunkingQiyang Li, Seohong Park, Sergey LevineICLR 2026 · 19 citations
- See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action ModelYixu Feng, Zinan Zhao, Yanxiang Ma, Chenghao Xia et al.ICML 2026 · 6 citations
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 1,720 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
Related papers
- Time Optimal Execution of Action Chunk Policies Beyond Demonstration SpeedSunwoo Kim, Jeongjun Kim, Joseph J LimICLR 2026
- TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic EnvironmentsZhiyu Huang, Yun Zhang, Johnson Liu, Rui Song et al.ICML 2026 · 11 citations
- Learning High-Frequency Continuous Action Chunks in Latent SpaceKunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao et al.ICML 2026
- Reflex: Real-Time Vision-Language-Action Control through Streaming InferenceYuanchun Guo, Bingyan LiuICML 2026
- Adaptive Action Chunking at Inference-time for Vision-Language-Action ModelsYuanchang Liang, Xiaobo Wang, Kai Wang, Shuo Wang et al.CVPR 2026 · 30 citations
