Find any Part in 3D
Ziqi Ma, Yisong Yue, Georgia Gkioxari
摘要
Why don't we have foundation models in 3D yet? A key limitation is data scarcity. For 3D object part segmentation, existing datasets are small in size and lack diversity. We show that it is possible to break this data barrier by building a data engine powered by foundation models. Our data engine automatically annotates any number of object parts: more unique part types than existing datasets combined. By training on our annotated data with a simple contrastive objective, we obtain an open-world model that generalizes to any part in any object based on any text query. Even when evaluated zero-shot, we outperform existing methods on the datasets they train on. We achieve improvement in mIoU and boost speed by to . Our scaling analysis confirms that this generalization stems from the data scale, which underscores the impact of our data engine. Finally, to advance general-category openworld 3D part segmentation, we release a benchmark covering a wide range of objects and parts. Project website: https://ziqi-ma.qithub.io/find3dsite/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht 等CVPR 2026 · 被引用 22 次
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang 等ICLR 2026 · 被引用 15 次
- DexVLG: Dexterous Vision-Language-Grasp Model at ScaleJiawei He, Danshi Li, Xinqiang Yu, Zekun Qi 等ICCV 2025 · 被引用 6 次
- PatchAlign3D: Local Feature Alignment for Dense 3D Shape UnderstandingSouhail Hadgi, Bingchen Gong, Ramana Sundararaman, Emery Pierson 等CVPR 2026 · 被引用 5 次
- Aligning Text, Images and 3D Structure Token-by-TokenAadarsh Sahoo, Vansh Tibrewal, Georgia GkioxariCVPR 2026 · 被引用 4 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai 等NeurIPS 2022 · 被引用 1,270 次
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu 等NeurIPS 2022 · 被引用 924 次
相关 Paper
- Point-SAM: Promptable 3D Segmentation Model for Point CloudsYuchen Zhou, Jiayuan Gu, Tung Yen Chiang, Fanbo Xiang 等ICLR 2025
- Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationJunha Lee, Chunghyun Park, Jaesung Choe, Yu-Chiang Frank Wang 等CVPR 2025
- InstructPart: Task-Oriented Part Segmentation with Instruction ReasoningZifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin 等ACL 2025 · 被引用 6 次
- Detect Anything 3D in the WildHanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun 等ICCV 2025 · 被引用 6 次
- SegVol: Universal and Interactive Volumetric Medical Image SegmentationYuxin Du, Fan Bai, Tiejun Huang, Bo ZhaoNeurIPS 2024 · 被引用 155 次
