Self-Supervised Viewpoint Learning From Image Collections
Siva Karthik Mustikovela, Varun Jampani, Shalini De Mello, Sifei Liu, Umar Iqbal, Carsten Rother, Jan Kautz
摘要
Training deep neural networks to estimate the viewpoint of objects requires large labeled training datasets. However, manually labeling viewpoints is notoriously hard, error-prone, and time-consuming. On the other hand, it is relatively easy to mine many unlabelled images of an object category from the internet, e.g., of cars or faces. We seek to answer the research question of whether such unlabeled collections of in-the-wild images can be successfully utilized to train viewpoint estimation networks for general object categories purely via self-supervision. Selfsupervision here refers to the fact that the only true supervisory signal that the network has is the input image itself. We propose a novel learning framework which incorporates an analysis-by-synthesis paradigm to reconstruct images in a viewpoint aware manner with a generative network, along with symmetry and adversarial constraints to successfully supervise our viewpoint estimation network. We show that our approach performs competitively to fullysupervised approaches for several object categories like human faces, cars, buses, and trains. Our work opens up further research in self-supervised viewpoint learning and serves as a robust baseline for it. We open-source our code at https://github.com/NVlabs/SSV . * Siva Karthik Mustikovela was an intern at NVIDIA during the project.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- GNeRF: GAN-based Neural Radiance Field without Posed CameraQuan Meng, Anpei Chen, Haimin Luo, Minye Wu 等ICCV 2021 · 被引用 222 次
- Convolutional Generation of Textured 3D MeshesDario Pavllo, Graham Spinks, Thomas Hofmann, Marie-Francine Moens 等NeurIPS 2020 · 被引用 71 次
- Self-Learning Transformations for Improving Gaze and Head RedirectionYufeng Zheng, Seonwook Park, Xucong Zhang, Shalini De Mello 等NeurIPS 2020 · 被引用 50 次
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
- Generative Modeling for Multi-task Visual LearningZhipeng Bao, Martial Hebert, Yu-Xiong WangICML 2022 · 被引用 18 次
它引用的顶会 Paper5
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 被引用 789 次
- Few-Shot Adaptive Gaze EstimationSeonwook Park, Shalini De Mello, Pavlo Molchanov, Umar Iqbal 等ICCV 2019 · 被引用 238 次
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 被引用 104 次
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt 等ICCV 2019 · 被引用 98 次
- Monocular Neural Image Based Rendering With Continuous View ControlJie Song, Xu Chen, Otmar HilligesICCV 2019 · 被引用 85 次
相关 Paper
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 被引用 8 次
- iVS-Net: Learning Human View Synthesis from Internet VideosJunting Dong, Qi Fang, Tianshuo Yang, Qing Shuai 等ICCV 2023 · 被引用 9 次
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D PoseAngtian Wang, Shenxiao Mei, Alan L. Yuille, Adam KortylewskiNeurIPS 2021 · 被引用 22 次
- Self-Supervised 3D Face Reconstruction via Conditional EstimationYandong Wen, Weiyang Liu, Bhiksha Raj, Rita SinghICCV 2021 · 被引用 24 次
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
