ViewNet: Unsupervised Viewpoint Estimation from Conditional Generation
Octave Mariotti, Oisin Mac Aodha, Hakan Bilen
Abstract
Understanding the 3D world without supervision is currently a major challenge in computer vision as the annotations required to supervise deep networks for tasks in this domain are expensive to obtain on a large scale. In this paper, we address the problem of unsupervised viewpoint estimation. We formulate this as a self-supervised learning task, where image reconstruction provides the supervision needed to predict the camera viewpoint. Specifically, we make use of pairs of images of the same object at training time, from unknown viewpoints, to self-supervise training by combining the viewpoint information from one image with the appearance information from the other. We demonstrate that using a perspective spatial transformer allows efficient viewpoint learning, outperforming existing unsupervised approaches on synthetic data, and obtains competitive results on the challenging PASCAL3D+ dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- NOPE: Novel Object Pose Estimation from a Single ImageVan Nguyen Nguyen, Thibault Groueix, Georgy Ponimatkin, Yinlin Hu et al.CVPR 2024 · 18 citations
- FisherMatch: Semi-Supervised Rotation Regression via Entropy-based FilteringYingda Yin, Yingcheng Cai, He Wang, Baoquan ChenCVPR 2022 · 16 citations
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 6 citations
Builds on9
- Texture Fields: Learning Texture Representations in Function SpaceMichael Oechsle, Lars M. Mescheder, Michael Niemeyer, Thilo Strauss et al.ICCV 2019 · 334 citations
- Escaping Plato's Cave: 3D Shape From Adversarial RenderingPhilipp Henzler, Niloy J. Mitra, Tobias RitschelICCV 2019 · 254 citations
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- Transformable Bottleneck NetworksKyle Olszewski, Sergey Tulyakov, Oliver J. Woodford, Hao Li et al.ICCV 2019 · 79 citations
- Single-Stage Semantic Segmentation From Image LabelsNikita Araslanov, Stefan RothCVPR 2020
Related papers
- Rayzer: a Self-Supervised Large View Synthesis ModelHanwen Jiang, Hao Tan, Peng Wang, Hai Jin et al.ICCV 2025 · 12 citations
- Self-Supervised Viewpoint Learning From Image CollectionsSiva Karthik Mustikovela, Varun Jampani, Shalini De Mello, Sifei Liu et al.CVPR 2020
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- From None to All: Self-Supervised 3D Reconstruction via Novel View SynthesisRanran Huang, Weixun Luo, Ye Mao, Krystian MikolajczykCVPR 2026 · 2 citations
- SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose EstimationVinkle Srivastav, Keqi Chen, Nicolas PadoyCVPR 2024 · 17 citations
