Self-Supervised Viewpoint Learning From Image Collections
Siva Karthik Mustikovela, Varun Jampani, Shalini De Mello, Sifei Liu, Umar Iqbal, Carsten Rother, Jan Kautz
Abstract
Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabelled images of an object category from the internet, e.g., of cars or faces. We seek to answer the research question of whether such unlabeled collections of in-the-wild images can be successfully utilized to train viewpoint estimation networks for general object categories purely via self-supervision. Selfsupervision here refers to the fact that the only true supervisory signal that the network has is the input image itself. We propose a novel learning framework which incorporates an analysis-by-synthesis paradigm to reconstruct images in a viewpoint aware manner with a generative network, along with symmetry and adversarial constraints to successfully supervise our viewpoint estimation network. We show that our approach performs competitively to fullysupervised approaches for several object categories like human faces, cars, buses, and trains. Our work opens up further research in self-supervised viewpoint learning and serves as a robust baseline for it. We open-source our code at https://github.com/NVlabs/SSV . * Siva Karthik Mustikovela was an intern at NVIDIA during the project.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- GNeRF: GAN-based Neural Radiance Field without Posed CameraQuan Meng, Anpei Chen, Haimin Luo, Minye Wu et al.ICCV 2021 · 222 citations
- Convolutional Generation of Textured 3D MeshesDario Pavllo, Graham Spinks, Thomas Hofmann, Marie-Francine Moens et al.NeurIPS 2020 · 71 citations
- Self-Learning Transformations for Improving Gaze and Head RedirectionYufeng Zheng, Seonwook Park, Xucong Zhang, Shalini De Mello et al.NeurIPS 2020 · 50 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Generative Modeling for Multi-task Visual LearningZhipeng Bao, Martial Hebert, Yu-Xiong WangICML 2022 · 18 citations
Builds on5
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Few-Shot Adaptive Gaze EstimationSeonwook Park, Shalini De Mello, Pavlo Molchanov, Umar Iqbal et al.ICCV 2019 · 238 citations
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 104 citations
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- Monocular Neural Image Based Rendering With Continuous View ControlJie Song, Xu Chen, Otmar HilligesICCV 2019 · 85 citations
Related papers
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 8 citations
- iVS-Net: Learning Human View Synthesis from Internet VideosJunting Dong, Qi Fang, Tianshuo Yang, Qing Shuai et al.ICCV 2023 · 9 citations
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D PoseAngtian Wang, Shenxiao Mei, Alan L. Yuille, Adam KortylewskiNeurIPS 2021 · 22 citations
- Self-Supervised 3D Face Reconstruction via Conditional EstimationYandong Wen, Weiyang Liu, Bhiksha Raj, Rita SinghICCV 2021 · 24 citations
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
