4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
Jiahui Zhang, Yurui Chen, Yueming Xu, Ze Huang, Yanpeng Zhou, Yu-Jie Yuan, Xinyue Cai, Guowei Huang, Xingyue Quan, Hang Xu, Li Zhang
摘要
Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset's action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution-an issue we refer to as coordinate system chaos and state chaos. This inconsistency significantly hampers pretraining efficiency. To address this, we propose 4D-VLA, a novel approach that effectively integrates 4D information into the input to mitigate these sources of chaos. Our model introduces depth and temporal information into visual features with sequential RGB-D inputs, aligning the coordinate systems of the robot and the scene. This alignment endows the model with strong spatiotemporal reasoning capabilities while minimizing training overhead. Additionally, we introduce memory bank sampling, a frame sampling strategy designed to extract informative frames from historical images, further improving effectiveness and efficiency. Experimental results demonstrate that our pretraining method and architectural components substantially enhance model performance. In both simulated and real-world experiments, our model achieves a significant increase in success rate over OpenVLA. To further assess spatial perception and generalization to novel views, we introduce MV-Bench, a multi-view simulation benchmark. Our model consistently outperforms existing methods, demonstrating stronger spatial understanding and adaptability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action ModelsLinqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong 等CVPR 2026 · 被引用 43 次
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu 等CVPR 2026 · 被引用 23 次
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA ManipulationJingjing Qian, Boyao Han, Chen Shi, Lei Xiao 等CVPR 2026 · 被引用 19 次
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan 等CVPR 2026 · 被引用 15 次
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human VideosYicheng Feng, Wanpeng Zhang, Ye Wang, Hao Luo 等CVPR 2026 · 被引用 14 次
它引用的顶会 Paper17
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
相关 Paper
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language ModelsPeiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma 等NeurIPS 2025 · 被引用 100 次
- UP-VLA: A Unified Understanding and Prediction Model for Embodied AgentJianke Zhang, Yanjiang Guo, Yucheng Hu, Xiaoyu Chen 等ICML 2025
- 4D Visual Pre-Training for Robot LearningChengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao 等ICCV 2025 · 被引用 1 次
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo 等CVPR 2026 · 被引用 6 次
- Spatial-Temporal Aware Visuomotor Diffusion Policy LearningZhenyang Liu, Yikai Wang, Kuanning Wang, Longfei Liang 等ICCV 2025 · 被引用 11 次
