DiffPoseNet: Direct Differentiable Camera Pose Estimation
Chethan M. Parameshwara, Gokul Hari, Cornelia Fermüller, Nitin J. Sanket, Yiannis Aloimonos
Abstract
Current deep neural network approaches for camera pose estimation rely on scene structure for 3D motion estimation, but this decreases the robustness and thereby makes cross-dataset generalization difficult. In contrast, classical approaches to structure from motion estimate 3D motion utilizing optical flow and then compute depth. Their accuracy, however, depends strongly on the quality of the optical flow. To avoid this issue, direct methods have been proposed, which separate 3D motion from depth estimation, but compute 3D motion using only image gradients in the form of normal flow. In this paper, we introduce a network NFlowNet, for normal flow estimation which is used to enforce robust and direct constraints. In particular, normal flow is used to estimate relative camera pose based on the cheirality (depth positivity) constraint. We achieve this by formulating the optimization problem as a differentiable cheirality layer, which allows for end-to-end learning of camera pose. We perform extensive qualitative and quantitative evaluation of the proposed DiffPoseNet's sensitivity to noise and its generalization across datasets. We compare our approach to existing state-of-the-art methods on KITTI, TartanAir, and TUM-RGBD datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 155ea5fe-36ce-4418-a8aa-4cbddead6d33Cited by top-tier papers6
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai et al.CVPR 2026 · 26 citations
- Learning Normal Flow Directly from EventsDehao Yuan, Levi Burner, Jiayi Wu, Minghui Liu et al.ICCV 2025 · 2 citations
- Robust Frame-to-Frame Camera Rotation Estimation in Crowded ScenesFabien Delattre, David Dirnfeld, Phat Nguyen, Stephen Scarano et al.ICCV 2023 · 2 citations
- Flow-Guided Online Stereo Rectification for Wide Baseline StereoAnush Kumar, Fahim Mannan, Omid Hosseini Jafari, Shile Li et al.CVPR 2024
- AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionLiuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo et al.ICCV 2025
Builds on3
- Attacking Optical FlowAnurag Ranjan, Joel Janai, Andreas Geiger, Michael J. BlackICCV 2019 · 93 citations
- MaskFlownet: Asymmetric Feature Matching With Learnable Occlusion MaskShengyu Zhao, Yilun Sheng, Yue Dong, Eric I-Chao Chang et al.CVPR 2020
- Towards Better Generalization: Joint Depth-Pose Learning Without PoseNetWang Zhao, Shaohui Liu, Yezhi Shu, Yong-Jin LiuCVPR 2020
Related papers
- Deep Two-View Structure-From-Motion RevisitedJianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield et al.CVPR 2021
- FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene FlowCameron Smith, Yilun Du, Ayush Tewari, Vincent SitzmannNeurIPS 2023 · 43 citations
- Consensus Learning with Deep Sets for Essential Matrix EstimationDror Moran, Yuval Margalit, Guy Trostianetsky, Fadi Khatib et al.NeurIPS 2024 · 4 citations
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 4 citations
- Polarimetric Relative Pose EstimationZhaopeng Cui, Viktor Larsson, Marc PollefeysICCV 2019 · 24 citations
