ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image Collections
Chun-Han Yao, Amit Raj, Wei-Chih Hung, Michael Rubinstein, Yuanzhen Li, Ming-Hsuan Yang, Varun Jampani
Abstract
Estimating 3D articulated shapes like animal bodies from monocular images is inherently challenging due to the ambiguities of camera viewpoint, pose, texture, lighting, etc. We propose ARTIC3D, a self-supervised framework to reconstruct per-instance 3D shapes from a sparse image collection in-the-wild. Specifically, ARTIC3D is built upon a skeleton-based surface representation and is further guided by 2D diffusion priors from Stable Diffusion. First, we enhance the input images with occlusions/truncation via 2D diffusion to obtain cleaner mask estimates and semantic features. Second, we perform diffusion-guided 3D optimization to estimate shape and texture that are of high-fidelity and faithful to input images. We also propose a novel technique to calculate more stable image-level gradients via diffusion models compared to existing alternatives. Finally, we produce realistic animations by fine-tuning the rendered shape and texture under rigid part transformations. Extensive evaluations on multiple existing datasets as well as newly introduced noisy web image collections with occlusions and truncation demonstrate that ARTIC3D outputs are more robust to noisy images, higher quality in terms of shape and texture details, and more realistic when animated. Project page: https://chhankyao.github.io/artic3d/ * Work done as a student researcher at Google. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f159d4e-b58d-4168-a903-8a831654ff31Cited by top-tier papers6
- VAREN: Very Accurate and Realistic Equine NetworkSilvia Zuffi, Ylva Mellbin, Ci Li, Markus Höschle et al.CVPR 2024 · 9 citations
- Primitive-Based 3D Human-Object Interaction Modelling and ProgrammingSiqi Liu, Yong-Lu Li, Zhou Fang, Xinpeng Liu et al.AAAI 2024 · 8 citations
- YouDream: Generating Anatomically Controllable Consistent Text-to-3D AnimalsSandeep Mishra, Oindrila Saha, Alan C. BovikNeurIPS 2024
- Learning the 3D Fauna of the WebZizhang Li, Dor Litvak, Ruining Li, Yunzhi Zhang et al.CVPR 2024
- MAS: Multi-view Ancestral Sampling for 3D Motion Generation Using 2D DiffusionRoy Kapon, Guy Tevet, Daniel Cohen-Or, Amit H. BermanoCVPR 2024
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part DiscoveryChun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein et al.NeurIPS 2022 · 83 citations
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
- Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic SamplingQi Sun, Can Wang, Jiaxiang Shang, Yingchun Liu et al.CVPR 2026 · 1 citation
- Online Adaptation for Consistent Mesh Reconstruction in the WildXueting Li, Sifei Liu, Shalini De Mello, Kihwan Kim et al.NeurIPS 2020 · 62 citations
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang et al.NeurIPS 2023 · 65 citations
