More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
Yuan Tang, Xu Han, Xianzhi Li, Qiao Yu, Jinfeng Xu, Yixue Hao, Long Hu, Min Chen
摘要
Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Exploring the Potential of Encoder-free Architectures in 3D LMMsYiwen Tang, Ziyu Guo, Zhuhao Wang, Renrui Zhang 等ICLR 2026 · 被引用 19 次
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene UnderstandingXiaohu Huang, Jingjing Wu, Qunyi Xie, Kai HanNeurIPS 2025 · 被引用 11 次
- Point Cloud as a Foreign Language for Multi-modal Large Language ModelSneha Paul, Zachary Patterson, Nizar BouguilaCVPR 2026 · 被引用 2 次
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 被引用 1 次
- Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud UnderstandingXianglong Jin, Zheng Wang, Rong Wang, Feiping NieAAAI 2026
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
相关 Paper
- MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D PriorsYuan Tang, Xu Han, Xianzhi Li, Qiao Yu 等ACM MM 2024 · 被引用 21 次
- GPT4Point: A Unified Framework for Point-Language Understanding and GenerationZhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu 等CVPR 2024 · 被引用 23 次
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo 等ICCV 2023 · 被引用 248 次
- ULIP: Learning a Unified Representation of Language, Images, and Point Clouds for 3D UnderstandingLe Xue, Mingfei Gao, Chen Xing, Roberto Martín-Martín 等CVPR 2023
- PointCLIP: Point Cloud Understanding by CLIPRenrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li 等CVPR 2022
