SAM 3D: 3Dfy Anything in Images
Xingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang, Alexander Sax, Hao Tang, Weiyao Wang, Michelle Guo, Thibaut Hardin, Xiang Li, Aohan Lin, Jia-Wei Liu
Abstract
We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual recognition cues from context play a larger role. We achieve this with a human- and model-in-the-loop pipeline for annotating object shape, texture, and pose, providing visually grounded 3D reconstruction data at unprecedented scale. We learn from this data in a modern, multi-stage training framework that combines synthetic pretraining with real-world alignment, breaking the 3D "data barrier". We obtain significant gains over recent work, with at least a win rate in human preference tests on real-world objects and scenes. We will release our code and model weights, an online demo, and a new challenging benchmark for in-the-wild 3D object reconstruction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c950efb-2503-4369-844e-4d9a72134f3bCited by top-tier papers16
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative ModelingElisabetta Fedele, Francis Engelmann, Ian Huang, Or Litany et al.ICLR 2026 · 11 citations
- Points-to-3D: Structure-Aware 3D Generation with Point Cloud PriorsJiatong Xia, Zicheng Duan, Anton van den Hengel, Lingqiao LiuCVPR 2026 · 6 citations
- Beyond Pixel Histories: World Models with Persistent 3D StateSamuel Garcin, Tom Walker, Steven McDonagh, Tim Pearce et al.ICML 2026 · 6 citations
- DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationYitong Yang, Yinglin Wang, Xuexin Liu, Jing Wang et al.ICML 2026 · 2 citations
- Affostruction: 3D Affordance Grounding with Generative ReconstructionChunghyun Park, Seunghyeon Lee, Minsu ChoCVPR 2026 · 1 citation
Builds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataWentao Wang, Hang Ye, Fangzhou Hong, Xue Yang et al.NeurIPS 2025 · 6 citations
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D SegmentationPanwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin et al.NeurIPS 2025
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D DataZhe Zhu, Le Wan, Rui Xu, Yiheng Zhang et al.ICLR 2026 · 15 citations
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao et al.NeurIPS 2025 · 8 citations
