PartGlot: Learning Shape Part Segmentation from Language Reference Games
Juil Koo, Ian Huang, Panos Achlioptas, Leonidas J. Guibas, Minhyuk Sung
Abstract
We introduce PartGlot, a neural framework and associated architectures for learning semantic part segmentation of 3D shape geometry, based solely on part referential language. We exploit the fact that linguistic descriptions of a shape can provide priors on the shape's parts - as natural language has evolved to reflect human perception of the compositional structure of objects, essential to their recognition and use. For training we use ShapeGlot's paired geometry /language data collected via a reference game where a speaker produces an utterance to differentiate a target shape from two distractors and the listener has to find the target based on this utterance [3]. Our network is designed to solve this target multi-modal recognition problem, by carefully incorporating a Transformer-based attention module so that the output attention can precisely highlight the semantic part or parts described in the language. Remarkably, the network operates without any direct supervision on the 3D geometry itself. Furthermore, we also demonstrate that the learned part information is generaliz-able to shape classes unseen during training. Our approach opens the possibility of learning 3D shape parts from language alone, without the need for large-scale part geometry annotations, thus facilitating annotation acquisition. The code is available at https://github.com/63days/PartGlot.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 199cce61-c4e5-4f31-bf8b-53c816406b53Cited by top-tier papers8
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 157 citations
- SALAD: Part-Level Latent Diffusion for 3D Shape Generation and ManipulationJuil Koo, Seungwoo Yoo, Minh Hieu Nguyen, Minhyuk SungICCV 2023 · 79 citations
- SATR: Zero-Shot Semantic Segmentation of 3D ShapesAhmed Abdelreheem, Ivan Skorokhodov, Maks Ovsjanikov, Peter WonkaICCV 2023 · 68 citations
- DiffFacto: Controllable Part-Based 3D Point Cloud Generation with Cross DiffusionGeorge Kiyohiro Nakayama, Mikaela Angelina Uy, Jiahui Huang, Shi-Min Hu et al.ICCV 2023 · 46 citations
- Iterative Superquadric Recomposition of 3D Objects from Multiple ViewsStephan Alaniz, Massimiliano Mancini, Zeynep AkataICCV 2023 · 21 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PoinTr: Diverse Point Cloud Completion with Geometry-Aware TransformersXumin Yu, Yongming Rao, Ziyi Wang, Zuyan Liu et al.ICCV 2021 · 592 citations
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna et al.ICCV 2019 · 427 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
Related papers
- Shapeglot: Learning Language for Shape DifferentiationPanos Achlioptas, Leonidas J. Guibas, Noah D. Goodman, Judy Fan et al.ICCV 2019 · 86 citations
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang et al.ICLR 2026 · 15 citations
- Part-X-MLLM: Part-aware 3D Multimodal Large Language ModelChunshi Wang, Junliang Ye, Yunhan Yang, YANG LI et al.ICLR 2026 · 6 citations
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying et al.CVPR 2026 · 1 citation
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to SentencesZhizhong Han, Chao Chen, Yu-Shen Liu, Matthias ZwickerACM MM 2020 · 38 citations
