PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
Qing Mao, Tianxin Huang, Yu Zhu, Jinqiu Sun, Yanning Zhang, Gim Hee Lee
Abstract
Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video interpolation and selecting key frames via a self-consistency score. However, the generated frames are often blurry due to small overlap inputs, and the selection strategies are slow and not explicitly aligned with pose estimation. To solve these cases, we propose Hybrid Video Generation (HVG) to synthesize clearer intermediate frames by coupling a video interpolation model with a pose-conditioned novel view synthesis model, where we also propose a Feature Matching Selector (FMS) based on feature correspondence to select intermediate frames appropriate for pose estimation from the synthesized results. Extensive experiments on Cambridge Landmarks, ScanNet, DL3DV-10K, and NAVI demonstrate that, compared to existing SOTA methods, PoseCrafter can obviously enhance the pose estimation performances, especially on examples with small or no overlap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
- StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2Ivan Skorokhodov, Sergey Tulyakov, Mohamed ElhoseinyCVPR 2022 · 167 citations
- Patch2Pix: Epipolar-Guided Pixel-Level CorrespondencesQunjie Zhou, Torsten Sattler, Laura Leal-TaixéCVPR 2021
Related papers
- ExPose: Reinforcing Video Generation Models for Extreme Pose EstimationYoungho Yoon, Wonjune Cho, Hyunho Ha, Sujung Kim et al.CVPR 2026
- Can Generative Video Models Help Pose Estimation?Ruojin Cai, Jason Y. Zhang, Philipp Henzler, Zhengqi Li et al.CVPR 2025
- SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsSongchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie et al.ICCV 2025 · 6 citations
- Extreme Relative Pose Network Under Hybrid RepresentationsZhenpei Yang, Siming Yan, Qixing HuangCVPR 2020
- Sparse-view Pose Estimation and Reconstruction via Analysis by Generative SynthesisQitao Zhao, Shubham TulsianiNeurIPS 2024 · 10 citations
