Lune

CVPR2026Top-tier venue

Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation

Yunlong Zhao, Xiaoheng Deng, Yichao Cao, Yi Chen, Xiangjian He, Shan You, Shuo Yang, Lei Fan, Fei Wang, Xiu Su

2026Year

Abstract

Diff Render +10.0 +10.0 +25.0 +20.0 +20.0 +20.0 Real-World Robot Manipulation Figure 1. 2D VLA models (left-top) leverage intuitive semantic perception from multi-view transformers but struggle with explicit 3D spatial reasoning. 3D VLA models (left-bottom) achieve precise spatial perception through point clouds and voxels but sacrifice visual interpretability. Our method bridges these paradigms by embedding spatial semantics into images via differentiable rendering.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ebf4cf83-ced1-413b-bce9-a84424d9b30d

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines