UprightNet: Geometry-Aware Camera Orientation Estimation From Single Images
Wenqi Xian, Zhengqi Li, Noah Snavely, Matthew Fisher, Jonathan Eisenmann, Eli Shechtman
Abstract
We introduce UprightNet, a learning-based approach for estimating 2DoF camera orientation from a single RGB image of an indoor scene. Unlike recent methods that leverage deep learning to perform black-box regression from image to orientation parameters, we propose an end-to-end framework that incorporates explicit geometric reasoning. In particular, we design a network that predicts two representations of scene geometry, in both the local camera and global reference coordinate systems, and solves for the camera orientation as the rotation that best aligns these two predictions via a differentiable least squares module. This network can be trained end-to-end, and can be supervised with both ground truth camera poses and intermediate representations of surface geometry. We evaluate UprightNet on the single-image camera orientation task on synthetic and real datasets, and show significant improvements over prior state-of-the-art approaches. * indicates equal contribution Weighted Least Square Y Z X Camera orientation Local camera surface geometry Global upright surface geometry Weights Input image
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6712e73a-1f75-408e-8ed4-1a5d01d9d282Cited by top-tier papers12
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe et al.ICCV 2021 · 246 citations
- CTRL-C: Camera calibration TRansformer with Line-ClassificationJinwoo Lee, Hyunsung Go, Hyunjoon Lee, Sunghyun Cho et al.ICCV 2021 · 52 citations
- Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised LearningIlwi Yun, Hyuk-Jae Lee, Chae-Eun RheeAAAI 2022 · 34 citations
- AnyCalib: On-Manifold Learning for Model-Agnostic Single-View Camera CalibrationJavier Tirado-Garín, Javier CiveraICCV 2025 · 9 citations
- Pop-Out Motion: 3D-Aware Image Deformation via Learning the Shape LaplacianJihyun Lee, Minhyuk Sung, Hyunjin Kim, Tae-Kyun KimCVPR 2022 · 3 citations
Builds on1
Related papers
- GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose EstimationGu Wang, Fabian Manhardt, Federico Tombari, Xiangyang JiCVPR 2021
- Back to the Feature: Learning Robust Camera Localization From Pixels To PosePaul-Edouard Sarlin, Ajaykumar Unagar, Måns Larsson, Hugo Germain et al.CVPR 2021
- Upright-Net: Learning Upright Orientation for 3D Point CloudXufang Pang, Feng Li, Ning Ding, Xiaopin ZhongCVPR 2022 · 11 citations
- Height and Uprightness Invariance for 3D Prediction From a Single ViewManel Baradad, Antonio TorralbaCVPR 2020
- Hierarchical Scene Coordinate Classification and Regression for Visual LocalizationXiaotian Li, Shuzhe Wang, Yi Zhao, Jakob Verbeek et al.CVPR 2020
