From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
Yiming Zhong, Yaoyu He, Zemin Yang, Pengfei Tian, Yifan Huang, Qingqiu Huang, Xinge Zhu, Yuexin Ma
Abstract
Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative VLA policies typically adopt a "Generation-from-Noise" paradigm, which disregards this disparity, leading to representation inefficiency and weak condition alignment during optimization. In this work, we propose ResVLA, an architecture that shifts the paradigm to "Refinement-from-Intent." Recognizing that robotic motion naturally decomposes into global intent and local dynamics, ResVLA utilizes spectral analysis to decouple control into a deterministic low-frequency anchor and a stochastic high-frequency residual. By anchoring the generative process on the predicted intent, our model focuses strictly on refining local dynamics via a residual diffusion bridge. Extensive simulation experiments show that ResVLA achieves competitive performance, strong robustness to language and robot embodiment perturbations, and faster convergence than standard generative baselines. ResVLA also demonstrates strong performance in real-world robot experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af8c05b2-38f7-4f45-85ff-3d69090d6750Builds on18
- Diffusion Schrödinger Bridge with Applications to Score-Based Generative ModelingValentin De Bortoli, James Thornton, Jeremy Heng, Arnaud DoucetNeurIPS 2021 · 811 citations
- I2SB: Image-to-Image Schrödinger BridgeGuan-Horng Liu, Arash Vahdat, De-An Huang, Evangelos A. Theodorou et al.ICML 2023 · 252 citations
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action ModelYihao Wang, Pengxiang Ding, Lingxiao Li, Can Cui et al.AAAI 2026 · 76 citations
Related papers
- Stable Language Guidance for Vision-Language-Action ModelsZhihao Zhan, Yuhao Chen, Jiaying Zhou, Qinhan Lyu et al.ACL 2026 · 7 citations
- LatentVLA: Taming Latent Space for Generalizable and Long-Horizon Bimanual ManipulationJunming WangAAAI 2026 · 1 citation
- SpikeVLA: Vision-Language-Action Models with Spiking Neural NetworksRuiqi Song, Dujun Nie, Siyu Teng, Baiyong Ding et al.ICML 2026 · 1 citation
- Scaling by Diversified Experience for Vision-Language-Action ModelsLeiyu Wang, Zhaofengnian Wang, Xueqi Li, Luoyi Fan et al.ICML 2026
- Libra-VLA: Achieving Learning Equilibrium via Asynchronous Coarse-to-Fine Dual-SystemYifei Wei, Linqing Zhong, Yi Liu, Yuxiang Lu et al.ACL 2026 · 1 citation
