Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
Xiuxiu Qi, Yu Yang, Jiannong Cao, Luyao Bai, Chongshan Fan, Chengtai Cao, Hongpeng Wang
Abstract
Language-Conditioned Manipulation (LCM) facilitates human-robot interaction via Behavioral Cloning (BC), which learns control policies from human demonstrations and serves as a cornerstone of embodied AI. Overcoming compounding errors in sequential action decisions remains a central challenge to improving BC performance. Existing approaches mitigate compounding errors through data augmentation, expressive representation, or temporal abstraction. However, they suffer from physical discontinuities and semantic-physical misalignment, leading to inaccurate action cloning and intermittent execution. In this paper, we present Continuous vision-language-action Co-Learning with Semantic-Physical Alignment (CCoL), a novel BC framework that ensures temporally consistent execution and fine-grained semantic grounding. It generates robust and smooth action execution trajectories through continuous co-learning across vision, language, and proprioceptive inputs (i.e., robot internal states). Meanwhile, we anchor language semantics to visuomotor representations by a bidirectional cross-attention to learn contextual information for action generation, successfully overcoming the problem of semantic-physical misalignment. Extensive experiments show that CCoL achieves an average 8.0% relative improvement across three simulation suites, with up to 19.2% relative gain in human-demonstrated bimanual insertion tasks. Real-world tests on a 7-DoF robot further confirm CCoL’s generalization under unseen and noisy object states.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8bc98ee-fe81-4b7c-8813-1dfe6b8bd979Builds on12
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- Language-Conditioned Imitation Learning for Robot Manipulation TasksSimon Stepputtis, Joseph Campbell, Mariano J. Phielipp, Stefan Lee et al.NeurIPS 2020 · 258 citations
- Neural Flows: Efficient Alternative to Neural ODEsMarin Bilos, Johanna Sommer, Syama Sundar Rangapuram, Tim Januschowski et al.NeurIPS 2021 · 151 citations
- Error Bounds of Imitating Policies and EnvironmentsTian Xu, Ziniu Li, Yang YuNeurIPS 2020 · 141 citations
- Is Behavior Cloning All You Need? Understanding Horizon in Imitation LearningDylan J. Foster, Adam Block, Dipendra MisraNeurIPS 2024 · 112 citations
Related papers
- Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied AgentsZhizhen Zhang, Lei Zhu, Zhen Fang, Zi Huang et al.NeurIPS 2025 · 5 citations
- Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAsJunhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji et al.ICML 2026
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic ManipulationHuayi Zhou, Kui JiaICLR 2026 · 3 citations
- Inner Speech as Behavior Guides: Steerable Imitation of Diverse Behaviors for Human-AI coordinationRakshit S. Trivedi, Kartik Sharma, David C. ParkesNeurIPS 2025 · 3 citations
- Language-Grounded Decoupled Action Representation for Robotic ManipulationWuDing Weng, Tongshu Wu, Liucheng Chen, Siyu xie et al.CVPR 2026 · 2 citations
