Shelf-Supervised Mesh Prediction in the Wild
Yufei Ye, Shubham Tulsiani, Abhinav Gupta
Abstract
We aim to infer 3D shape and pose of object from a single image and propose a learning-based approach that can train from unstructured image collections, supervised by only segmentation outputs from off-the-shelf recognition systems (i.e. ‘shelf-supervised’). We first infer a volumetric representation in a canonical frame, along with the camera pose. We enforce the representation geometrically consistent with both appearance and masks, and also that the synthesized novel views are indistinguishable from image collections. The coarse volumetric prediction is then converted to a mesh-based representation, which is further refined in the predicted camera frame. These two steps allow both shape-pose factorization from image collections and per-instance reconstruction in finer details. We examine the method on both synthetic and the real-world datasets and demonstrate its scalability on 50 categories in the wild, an order of magnitude more classes than existing works.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5349c696-67ff-41d9-8cb0-2cb502898595Cited by top-tier papers35
- EpiGRAF: Rethinking training of 3D GANsIvan Skorokhodov, Sergey Tulyakov, Yiqun Wang, Peter WonkaNeurIPS 2022 · 145 citations
- BANMo: Building Animatable 3D Neural Models from Many Casual VideosGengshan Yang, Minh Vo, Natalia Neverova, Deva Ramanan et al.CVPR 2022 · 113 citations
- ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape ReconstructionGengshan Yang, Deqing Sun, Varun Jampani, Daniel Vlasic et al.NeurIPS 2021 · 103 citations
- What's in your hands? 3D Reconstruction of Generic Objects in HandsYufei Ye, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 69 citations
- PPR: Physically Plausible Reconstruction from Monocular VideosGengshan Yang, Shuo Yang, John Z. Zhang, Zachary Manchester et al.ICCV 2023 · 41 citations
Builds on8
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Escaping Plato's Cave: 3D Shape From Adversarial RenderingPhilipp Henzler, Niloy J. Mitra, Tobias RitschelICCV 2019 · 254 citations
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 104 citations
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 72 citations
Related papers
- Pretrain, Self-train, Distill: A simple recipe for Supersizing 3D ReconstructionKalyan Vasudev Alwala, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 23 citations
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the WildShangzhe Wu, Christian Rupprecht, Andrea VedaldiCVPR 2020
- Topologically-Aware Deformation Fields for Single-View 3D ReconstructionShivam Duggal, Deepak PathakCVPR 2022 · 30 citations
- C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From MotionDavid Novotný, Nikhila Ravi, Benjamin Graham, Natalia Neverova et al.ICCV 2019 · 126 citations
