Test-Time Optimization of 3D Point Cloud LLM via Manifold-Aware In-Context Guidance and Refinement
Tiankai Chen, Nanqing Liu, Li Yang, Xulei Yang, Tianrui Li, Xun Xu
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in textual and 2D visual reasoning, yet their ability to understand and reason over 3D data remains limited. The issues become more challenging for understanding standalone 3D point cloud due to the high interclass confusion. In this work, we propose Point-Graph LLM (PGLLM), a framework that enables more effective 3D point cloud understanding by integrating in-context prompting and score refinement at test-time, respecting supporting data manifold. Our method first employs a pre-trained point cloud encoder which are used to construct a graph where edges encode visual similarity. Each support point cloud sample is converted to a textual caption via pre-trained PointLLM. For a test query, the graph is used to retrieve relevant neighbors whose captions serve as contextual demonstrations for a second stage LLM for final reasoning, a process we term in-context guidance. Furthermore, we introduce a confidence score refinement mechanism based on label propagation to enhance the reliability of LLM predictions for classification and out-of-distribution (OOD) detection tasks. All the above optimizations are carried out fully at test-time. Extensive experiments across diverse 3D datasets and tasks demonstrate that PGLLM consistently improves accuracy and robustness over prior baselines with very almost no additional computation cost, showcasing a promising direction toward native 3D reasoning with MLLMs. The code is available on https://github.com/handsome999KK/PGLLM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8bc3f92-25f8-4d18-a6c7-15a0ad50f9eaBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
Related papers
- PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-ThoughtChaoqi Chen, Qile Xu, Wenjun Zhou, Hui HuangSIGGRAPH 2026
- Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationTiankai Chen, Yushu Li, Adam Goodge, Fei Teng et al.ICCV 2025
- LL3DA: Visual Interactive Instruction Tuning for Omni-3D Understanding, Reasoning, and PlanningSijin Chen, Xin Chen, Chi Zhang, Mingsheng Li et al.CVPR 2024
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 130 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
