Find any Part in 3D
Ziqi Ma, Yisong Yue, Georgia Gkioxari
Abstract
Why don't we have foundation models in 3D yet? A key limitation is data scarcity. For 3D object part segmentation, existing datasets are small in size and lack diversity. We show that it is possible to break this data barrier by building a data engine powered by foundation models. Our data engine automatically annotates any number of object parts: more unique part types than existing datasets combined. By training on our annotated data with a simple contrastive objective, we obtain an open-world model that generalizes to any part in any object based on any text query. Even when evaluated zero-shot, we outperform existing methods on the datasets they train on. We achieve improvement in mIoU and boost speed by to . Our scaling analysis confirms that this generalization stems from the data scale, which underscores the impact of our data engine. Finally, to advance general-category openworld 3D part segmentation, we release a benchmark covering a wide range of objects and parts. Project website: https://ziqi-ma.qithub.io/find3dsite/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht et al.CVPR 2026 · 22 citations
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang et al.ICLR 2026 · 15 citations
- DexVLG: Dexterous Vision-Language-Grasp Model at ScaleJiawei He, Danshi Li, Xinqiang Yu, Zekun Qi et al.ICCV 2025 · 6 citations
- PatchAlign3D: Local Feature Alignment for Dense 3D Shape UnderstandingSouhail Hadgi, Bingchen Gong, Ramana Sundararaman, Emery Pierson et al.CVPR 2026 · 5 citations
- Aligning Text, Images and 3D Structure Token-by-TokenAadarsh Sahoo, Vansh Tibrewal, Georgia GkioxariCVPR 2026 · 4 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
Related papers
- Point-SAM: Promptable 3D Segmentation Model for Point CloudsYuchen Zhou, Jiayuan Gu, Tung Yen Chiang, Fanbo Xiang et al.ICLR 2025
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationJunha Lee, Chunghyun Park, Jaesung Choe, Yu-Chiang Frank Wang et al.CVPR 2025
- InstructPart: Task-Oriented Part Segmentation with Instruction ReasoningZifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin et al.ACL 2025 · 6 citations
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun et al.ICCV 2025 · 6 citations
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 155 citations
