Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
Xiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li, Sanglu Lu
Abstract
Vision-Language-Action (VLA) models built upon Chain-of-Thought (CoT) have achieved remarkable success in advancing general-purpose robotic agents, owing to its significant perceptual comprehension. Recently, since text-only CoT struggles to adequately capture scene details in complex spatial environments, a highly promising strategy involves leveraging visual priors to guide robotic action generation. Nevertheless, these strategies face two inherent challenges: (i) a modality gap between visual observations and low-level actions, and (ii) unstable training due to competing objectives between visual prediction and action generation. To address these challenges, we propose a Vision-Integrated Trajectory Alignment (VITA) framework that learns a shared discrete latent space for vision and action, enabling joint modeling of perception and motor control. VITA introduces a implicit visual CoT: autoregressively generated tokens is simultaneously decoded into future frames predictions and robot actions, thereby internalizing visual dynamics as an inductive bias for motion planning. Extensive experiments on simulated and real-world environments demonstrate state-of-the-art performance. VITA improves 14.5%, 9.6% and 12.1% over existing baselines on CALVIN, LIBERO and SimplerEnv. Furthermore, VITA attains an average success rate of 80.5% across six real-world tasks, demonstrating its potential as a generalist robotic manipulation model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8086437-0544-4190-a7c3-180df903ee1eCited by top-tier papers3
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
- Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action ModelsShuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li et al.ICML 2026 · 15 citations
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma et al.ICML 2026 · 2 citations
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action ModelJinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu et al.ICLR 2026 · 335 citations
Related papers
- CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsQingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu et al.CVPR 2025
- Unified Vision-Language-Action ModelYuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang et al.ICLR 2026 · 144 citations
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong et al.CVPR 2026 · 43 citations
- TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action ModelXiang Li, Ya-Li Li, Yuan Wang, Huaqiang Wang et al.AAAI 2026
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
