Novel Object Viewpoint Estimation Through Reconstruction Alignment
Mohamed El Banani, Jason J. Corso, David F. Fouhey
Abstract
The goal of this paper is to estimate the viewpoint for a novel object. Standard viewpoint estimation approaches generally fail on this task due to their reliance on a 3D model for alignment or large amounts of class-specific training data and their corresponding canonical pose. We overcome those limitations by learning a reconstruct and align approach. Our key insight is that although we do not have an explicit 3D model or a predefined canonical pose, we can still learn to estimate the object's shape in the viewer's frame and then use an image to provide our reference model or canonical pose. In particular, we propose learning two networks: the first maps images to a 3D geometry-aware feature bottleneck and is trained via an image-to-image translation loss; the second learns whether two instances of features are aligned. At test time, our model finds the relative transformation that best aligns the bottleneck features of our test image to a reference image. We evaluate our method on novel object viewpoint estimation by generalizing across different datasets, analyzing the impact of our different modules, and providing a qualitative analysis of the learned features to identify what representations are being learnt for alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efd5c328-0255-4b7e-b1af-8b0b3bde9ae0Cited by top-tier papers4
- Planar Surface Reconstruction from Sparse ViewsLinyi Jin, Shengyi Qian, Andrew Owens, David F. FouheyICCV 2021 · 51 citations
- E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-trainingQitao Zhao, Hao Tan, Qianqian Wang, Sai Bi et al.CVPR 2026 · 24 citations
- Virtual Correspondence: Humans as a Cue for Extreme-View GeometryWei-Chiu Ma, Anqi Joyce Yang, Shenlong Wang, Raquel Urtasun et al.CVPR 2022 · 20 citations
- UnsupervisedR&R: Unsupervised Point Cloud Registration via Differentiable RenderingMohamed El Banani, Luya Gao, Justin JohnsonCVPR 2021
Related papers
- Unsupervised Learning of Category-Level 3D Pose from Object-Centric VideosLeonhard Sommer, Artur Jesslen, Eddy Ilg, Adam KortylewskiCVPR 2024
- MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training SupervisionRuicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang et al.CVPR 2025
- Self-Supervised Learning of Interpretable Keypoints From Unlabelled VideosTomas Jakab, Ankush Gupta, Hakan Bilen, Andrea VedaldiCVPR 2020
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 8 citations
- Reconstruct Locally, Localize Globally: A Model Free Method for Object Pose EstimationMing Cai, Ian ReidCVPR 2020
