Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation
Zhaoyang Liu, Mokai Pan, Zhongyi Wang, Kaizhen Zhu, Haotao Lu, Haipeng Zhang, Jingya Wang, Ye Shi
Abstract
Imitation learning with diffusion models has advanced robotic control by capturing the multimodal action distributions. However, existing methods typically treat observations only as highlevel conditions to the denoising network, rather than integrating them into the stochastic dynamics of the diffusion process itself. As a result, the sampling is forced to begin from random noise, weakening the coupling between perception and control and often yielding suboptimal performance. We propose BridgePolicy, a generative visuomotor policy that directly integrates observations into the stochastic dynamics via a diffusion-bridge formulation. By constructing an observationinformed trajectory, BridgePolicy enables sampling to start from a rich and informative prior rather than random noise, substantially improving precision and reliability in control. A key difficulty is that diffusion bridge normally connects distributions of matched dimensionality, while robotic observations are heterogeneous and not naturally aligned with actions. To overcome this, we introduce a semantic aligner to unify the visual and state inputs and align the observations with action representations, making diffusion bridge applicable to heterogeneous robot data. Extensive experiments across 52 simulation tasks on three benchmarks and 5 real-world tasks demonstrate that BridgePolicy consistently outperforms state-of-the-art generative policies. Our code is available at https://jianghcsr.github. io/BridgePolicy_page/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dda000a-27ce-421b-8eda-546e982a3b88Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Prior Does Matter: Visual Navigation via Denoising Diffusion Bridge ModelsHao Ren, Yiming Zeng, Zetong Bi, Zhaoliang Wan et al.CVPR 2025
- Prediction with Action: Visual Policy Learning via Joint Denoising ProcessYanjiang Guo, Yucheng Hu, Jianke Zhang, Yen-Jen Wang et al.NeurIPS 2024 · 93 citations
- STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency PredictionJinhao Li, Yuxuan Cong, Yingqiao Wang, Hao Xia et al.ICML 2026 · 5 citations
- DynBridge: Bridging Imagination and Control through Interaction Dynamics for Robot ManipulationAlex Wang, Zhiwei Dong, Qicheng Bai, Chenshi Zhang et al.CVPR 2026
- Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive PoliciesJing Wang, Weiting Peng, Jing Tang, Zeyu Gong et al.NeurIPS 2025
