Lune

CVPR2026Top-tier venue

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation

Wei Li, Jizhihui Liu, Yixing Li, Junwen Tong, Rui Shao, Liqiang Nie

2026Year
8Citations
3Top-tier citations

Abstract

Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions but exhibit notable limitations in spatiotemporal perception and reasoning: 1) spatial representations often rely on additional sensors, introducing substantial computational overhead; 2) visual reasoning is typically limited to future-frame prediction, lacking alignment with the instruction-grounded scene and thus compromising spatiotemporal consistency. To address these challenges, we propose ConsisVLA-4D , a unified and efficient framework that enhances spatiotemporal consistency in 3D-Perception and 4D-Reasoning. Specifically, we design: 1) CV-Aligner , which ensures C ross- V iew object semantic consistency via filtering instruction-relevant regions and aligning object identities across multiple viewpoints; 2) CO-Fuser , which guarantees C ross- O bject spatial geometric consistency by eliminating spatial relation ambiguities between objects across views using compact latent representations. Building upon these, we introduce 3) CS-Thinker to achieve C ross- S cene spatiotemporal consistency as actions unfold. It learns implicit knowledge of local dynamics from object-semantic tokens of CV-Aligner and global depth from geometric tokens of CO-Fuser, thereby enhancing efficient visual reasoning under scene variations. Extensive experiments demonstrate that, benefiting from its efficient spatiotemporal consistency design, ConsisVLA-4D achieves 21.6% and 41.5% performance improvements, along with 2.3× and 2.4× inference speedups compared to OpenVLA on the LIBERO benchmark and real-world platforms, respectively.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 8474f6a9-5481-40f7-8082-e920f63a2a72

Cited by top-tier papers3

Ask how each one uses it

Builds on42

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines