LLaVA-3D: A Simple Yet Effective Pathway to Empowering LMMs with 3D Capabilities
Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, Xihui Liu
Abstract
Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the development of LMMs with 3D scene understanding capabilities has been hindered by the lack of large-scale 3D vision-language datasets and powerful 3D encoders. In this paper, we introduce a simple yet effective framework called LLaVA-3D. Leveraging the strong 2D visual understanding priors from LLaVA, our LLaVA-3D efficiently adapts LLaVA for 3D scene understanding without compromising 2D understanding capabilities. To achieve this, we utilize the 3D position embeddings to enhance the 2D CLIP Patches with 3D spatial context information and construct 3D patches. By integrating the 3D position embeddings into 2D LMMs and employing joint 2D and 3D vision-language instruction tuning, we establish a unified architecture for both 2D visual understanding and 3D scene understanding. In contrast to previous 3D LMMs, LLaVA-3D supports decoding accurate 3D spatial perception outputs, e.g., 3D bounding boxes, directly from these 3D patches, without relying on the time-consuming off-the-shelf 3D segmentors. Experimental results show that LLaVA-3D converges faster than existing 3D LMMs when trained on 3D vision-language datasets. Moreover, LLaVA-3D not only achieves state-of-the-art performance across various 3D tasks but also maintains comparable 2D visual understanding and vision-language conversation capabilities with LLaVA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9f2bc83-b494-4472-b8ce-82b3cd730c61Cited by top-tier papers46
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 130 citations
- Cameras as Relative Positional EncodingRuilong Li, Brent Yi, Junchen Liu, Hang Gao et al.NeurIPS 2025 · 113 citations
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng et al.AAAI 2026 · 35 citations
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous DrivingPeizheng Li, Zhenghao Zhang, David Holtz, Hang Yu et al.CVPR 2026 · 32 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene UnderstandingHanyu Zhou, Gim Hee LeeICLR 2026 · 20 citations
- LLaVA³: Representing 3D Scenes Like a Cubist Painter to Boost 3D Scene Understanding of VLMsDoriand Petit, Steve Bourgeois, Vincent Gay-Bellile, Florian Chabot et al.AAAI 2026
- Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene UnderstandingDuo Zheng, Shijia Huang, Liwei WangCVPR 2025
- LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR UnderstandingSenqiao Yang, Jiaming Liu, Renrui Zhang, Mingjie Pan et al.AAAI 2025 · 17 citations
- 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint TransformerJiajun Deng, Tianyu He, Li Jiang, Tianyu Wang et al.CVPR 2025
