ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
Zhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu
2025Year
6Citations
12Top-tier citations
Abstract
Visual Grounding Q: Provide the bounding box coordinate of the police vehicle. A: [0.26, 0.56, 0.44, 0.71] Image Captioning Q: Provide a one-sentence caption for the image. A: A vintage-style street clock stands prominently at a city intersection, with a historic brick building in the background and several cars, including a police car, navigating the crosswalk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b73f50c-67f1-4de9-a6c5-3349679ad97dCited by top-tier papers12
- CogVLA: Cognition-Aligned Vision-Language-Action Models via Instruction-Driven Routing & SparsificationWei Li, Renshan Zhang, Rui Shao, Jie He et al.NeurIPS 2025 · 87 citations
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich ManipulationYang Li, Zhaxizhuoma, Hongru Jiang, Junjie Xia et al.CVPR 2026 · 31 citations
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsBorong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang et al.ICML 2026 · 25 citations
- VLANeXt: Recipes for Building Strong VLA ModelsXiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang et al.ICML 2026 · 10 citations
Builds on16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
Related papers
- Leveraging Panoptic Scene Graph for Evaluating Fine-Grained Text-to-Image GenerationXueqing Deng, Linjie Yang, Qihang Yu, Chenglin Yang et al.ICCV 2025 · 2 citations
- It's About Time: Analog Clock Reading in the WildCharig Yang, Weidi Xie, Andrew ZissermanCVPR 2022 · 16 citations
- CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video GenerationQinghe Wang, Yawen Luo, Xiaoyu Shi, Xu Jia et al.SIGGRAPH 2025 · 13 citations
- MAVIS: Map-Aided Vehicular ISAC SynchronizationNikhil K. Nataraja, Kamran Ali, Fan Bai, Yuning Zhang et al.INFOCOM 2026
- Crab: A Unified Audio-Visual Scene Understanding Model with Explicit CooperationHenghui Du, Guangyao Li, Chang Zhou, Chunjie Zhang et al.CVPR 2025
