Conditioning Matters: Training Diffusion Policies is Faster Than You Think
Zibin Dong, Yicheng Liu, Yinchuan Li, Hang Zhao, Jianye Hao
摘要
Diffusion policies have emerged as a mainstream paradigm for building visionlanguage-action (VLA) models. Although they demonstrate strong robot control capabilities, their training efficiency remains suboptimal. In this work, we identify a fundamental challenge in conditional diffusion policy training: when generative conditions are hard to distinguish, the training objective degenerates into modeling the marginal action distribution, a phenomenon we term loss collapse. To overcome this, we propose Cocos, a simple yet general solution that modifies the source distribution in the conditional flow matching to be condition-dependent. By anchoring the source distribution around semantics extracted from condition inputs, Cocos encourages stronger condition integration and prevents the loss collapse. We provide theoretical justification and extensive empirical results across simulation and real-world benchmarks. Our method achieves faster convergence and higher success rates than existing approaches, matching the performance of large-scale pre-trained VLAs using significantly fewer gradient steps and parameters. Cocos is lightweight, easy to implement, and compatible with diverse policy architectures, offering a general-purpose improvement to diffusion policy training.
Figure 1: Fusing generative condition into the source distribution greatly simplifies diffusion policy training. Diffusion policy trained with our method achieves π0 performance on the LIBERO benchmarks with only 30K gradient steps, which is 2.14x faster than the vanilla model. We also show the cosine similarity and the norm scale change between the policy hidden states before and after injecting condition information, demonstrating that our method fundamentally compels the policy network to utilize condition information, rather than simply embedding conditions into the source distribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency PredictionJinhao Li, Yuxuan Cong, Yingqiao Wang, Hao Xia 等ICML 2026 · 被引用 5 次
- FASTer: Toward Powerful and Efficient Autoregressive Vision-Language-Action Models with Learnable Action Tokenizer and Block-wise DecodingYicheng Liu, Shiduo Zhang, Zibin Dong, Baijun Ye 等ICLR 2026
- From Noise to Intent: Anchoring Generative VLA Policies with Residual BridgesYiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian 等ICML 2026
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data AugmentationChenyu Hui, Xiaodi Huang, Siyu Xu, Yunke Wang 等ICML 2026 · 被引用 2 次
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic AlignmentTao Lin, Yilei Zhong, Yuxin Du, Jingjing Zhang 等CVPR 2026 · 被引用 47 次
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An 等ICLR 2026 · 被引用 216 次
- Exploring Conditions for Diffusion Models in Robotic ControlHeeseong Shin, Byeongho Heo, Dongyoon Han, Seungryong Kim 等CVPR 2026
- VITA: Vision-to-Action Flow Matching PolicyDechen Gao, BOQI ZHAO, Andrew Lee, Ian Chuang 等ICLR 2026 · 被引用 27 次
