Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description
Mahmoud Ahmed, Junjie Fei, Jian Ding, Eslam Mohamed Bakr, Mohamed Elhoseiny
Abstract
In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, partaware segmentation grounding and detailed explanation of 3D objects. Existing 3D datasets largely focus on either vision-only part segmentation or vision-language scene segmentation, lacking the fine-grained multimodal segmentation needed for robotic navigation and interaction in real-world environments. To address this gap, we present the 3DCoMPaT Grounded Instructions (3DCoMPaT-GrIn) Dataset, a comprehensive resource that pairs rich point cloud descriptions with corresponding part-level segmentation masks. This dataset encompasses extensive samples designed for both PaPGD and fine-grained singlepart grounding tasks. To tackle the inherent challenges of grounding objects and generating grounded descriptions at the part level, we propose Kestrel, a part-aware 3D multimodal large language model that integrates an advanced language model for nuanced language comprehension with multi-level point feature propagation and query refinement mechanism to enhance spatial reasoning at the part level. The extensive experiments demonstrate that Kestrel effectively bridges the gap between part-aware language understanding and 3D segmentation grounding, paving the way for more robust and interpretable 3D object comprehension that meets the demands of real-world robotic applications. Project page: https://feielysia.github. io/Kestrel.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6fe7249-18dc-40e0-927f-ba309b1ae90dCited by top-tier papers2
- BANG: Dividing 3D Assets via Generative Exploded DynamicsLongwen Zhang, Qixuan Zhang, Haoran Jiang, Yinuo Bai et al.SIGGRAPH 2025 · 5 citations
- SimArt: Decomposing Monolithic Meshes into Sim-ready Articulated Assets via MLLMChuanrui Zhang, Minghan Qin, Yuang Wang, Baifeng Xie et al.SIGGRAPH 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang et al.ICCV 2025 · 1 citation
- InstructPart: Task-Oriented Part Segmentation with Instruction ReasoningZifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin et al.ACL 2025 · 6 citations
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang et al.ICCV 2021 · 115 citations
- Semantic Guided Part Relation-aware Network for Point Cloud CompletionZhensheng Zhou, Jianqing Liang, Jiye Liang, Zijin Du et al.AAAI 2026
- Talking Points: Describing and Localizing PixelsMatan Rusanovsky, Shimon Malnick, Shai AvidanICLR 2026
