MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
Yuan Tang, Xu Han, Xianzhi Li, Qiao Yu, Yixue Hao, Long Hu, Min Chen
Abstract
Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also integrate point clouds into LLMs. However, directly aligning point clouds with LLM requires expensive training costs, typically in hundreds of GPU-hours on A100, which hinders the development of 3D-LLMs. In this paper, we introduce MiniGPT-3D, an efficient and powerful 3D-LLM that achieves multiple SOTA results while training for only 27 hours on one RTX 3090. Specifically, we propose to align 3D point clouds with LLMs using 2D priors from 2D-LLMs, which can leverage the similarity between 2D and 3D visual information. We introduce a novel four-stage training strategy for modality alignment in a cascaded way, and a mixture of query experts module to adaptively aggregate features with high efficiency. Moreover, we utilize parameter-efficient fine-tuning methods LoRA and Norm fine-tuning, resulting in only 47.8M learnable parameters, which is up to 260x fewer than existing methods. Extensive experiments show that MiniGPT-3D achieves SOTA on 3D object classification and captioning tasks, with significantly cheaper training costs. Notably, MiniGPT-3D gains an 8.12 increase on GPT-4 evaluation score for the challenging object captioning task compared to ShapeLLM-13B, while the latter costs 160 total GPU-hours on 8 A800. We are the first to explore the efficient 3D-LLM, offering new insights to the community. Code and weights are available at https://github.com/TangYuan96/MiniGPT-3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f237b092-8d0f-4615-91a1-9ef70d736e29Cited by top-tier papers24
- 3D Aware Region Prompted Vision Language ModelAn-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu et al.ICLR 2026 · 30 citations
- Exploring the Potential of Encoder-free Architectures in 3D LMMsYiwen Tang, Ziyu Guo, Zhuhao Wang, Renrui Zhang et al.ICLR 2026 · 19 citations
- Generated and Pseudo Content guided Prototype Refinement for Few-shot Point Cloud SegmentationLili Wei, Congyan Lang, Ziyi Chen, Tao Wang et al.NeurIPS 2024 · 12 citations
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene UnderstandingXiaohu Huang, Jingjing Wu, Qunyi Xie, Kai HanNeurIPS 2025 · 11 citations
- Kestrel: 3D Multimodal LLM for Part-Aware Grounded DescriptionMahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr et al.ICCV 2025 · 9 citations
Builds on24
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 1 citation
- More Text, Less Point: Towards 3D Data-Efficient Point-Language UnderstandingYuan Tang, Xu Han, Xianzhi Li, Qiao Yu et al.AAAI 2025 · 6 citations
- GPT4Point: A Unified Framework for Point-Language Understanding and GenerationZhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu et al.CVPR 2024 · 23 citations
- CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain AdaptationMainak Singha, Sarthak Mehrotra, Paolo Casari, Subhasis Chaudhuri et al.CVPR 2026 · 2 citations
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang et al.CVPR 2026 · 179 citations
