SAM 3D: 3Dfy Anything in Images
Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jia-Wei Liu
摘要
We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual recognition cues from context play a larger role. We achieve this with a human- and model-in-the-loop pipeline for annotating object shape, texture, and pose, providing visually grounded 3D reconstruction data at unprecedented scale. We learn from this data in a modern, multi-stage training framework that combines synthetic pretraining with real-world alignment, breaking the 3D "data barrier". We obtain significant gains over recent work, with at least a win rate in human preference tests on real-world objects and scenes. We will release our code and model weights, an online demo, and a new challenging benchmark for in-the-wild 3D object reconstruction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative ModelingElisabetta Fedele, Francis Engelmann, Ian Huang, Or Litany 等ICLR 2026 · 被引用 11 次
- Points-to-3D: Structure-Aware 3D Generation with Point Cloud PriorsJiatong Xia, Zicheng Duan, Anton van den Hengel, Lingqiao LiuCVPR 2026 · 被引用 6 次
- Beyond Pixel Histories: World Models with Persistent 3D StateSamuel Garcin, Tom Walker, Steven McDonagh, Tim Pearce 等ICML 2026 · 被引用 6 次
- DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationYitong Yang, Yinglin Wang, Xuexin Liu, Jing Wang 等ICML 2026 · 被引用 2 次
- Affostruction: 3D Affordance Grounding with Generative ReconstructionChunghyun Park, Seunghyeon Lee, Minsu ChoCVPR 2026 · 被引用 1 次
它引用的顶会 Paper44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan 等CVPR 2026 · 被引用 81 次
- GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataWentao Wang, Hang Ye, Fangzhou Hong, Xue Yang 等NeurIPS 2025 · 被引用 6 次
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D SegmentationPanwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin 等NeurIPS 2025
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang 等ICLR 2026 · 被引用 15 次
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao 等NeurIPS 2025 · 被引用 8 次
