H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
Heng Jia, Linchao Zhu, Na Zhao
Abstract
Despite recent advances in feed-forward 3D Gaussian Splatting, generalizable 3D reconstruction remains challenging, particularly in multi-view correspondence modeling. Existing approaches face a fundamental trade-off: explicit methods achieve geometric precision but struggle with ambiguous regions, while implicit methods provide robustness but suffer from slow convergence. We present H3R, a hybrid framework that addresses this limitation by integrating volumetric latent fusion with attention-based feature aggregation. Our framework consists of two complementary components: an efficient latent volume that enforces geometric consistency through epipolar constraints, and a camera-aware Transformer that leverages Plücker coordinates for adaptive correspondence refinement. By integrating both paradigms, our approach enhances generalization while converging 2 faster than existing methods. Furthermore, we show that spatial-aligned foundation models (e.g., SD-VAE) substantially outperform semantic-aligned models (e.g., DINOv2), resolving the mismatch between semantic representations and spatial reconstruction requirements. Our method supports variable-number and high-resolution input views while demonstrating robust cross-dataset generalization. Extensive experiments show that our method achieves state-of-the-art performance across multiple benchmarks, with significant PSNR improvements of 0.59 dB, 1.06 dB, and 0.22 dB on the RealEstate10K, ACID, and DTU datasets, respectively. Code is available at https://github.com/JiaHeng-DLUT/H3R.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa39d442-bc39-4ed1-94cd-cb44c5698528Cited by top-tier papers1
Ask how each one uses itBuilds on46
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation ModelsYifan Liu, Keyu Fan, Weihao Yu, Chenxin Li et al.CVPR 2025
- DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View ConsistencyTengfei Xiao, Yue Wu, Zhigang Gao, Yongzhe Yuan et al.AAAI 2026
- GSRecon: Efficient Generalizable Gaussian Splatting for Surface Reconstruction from Sparse ViewsHang Yang, Le Hui, Jianjun Qian, Jin Xie et al.ICCV 2025 · 1 citation
- Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian SplattingYiming Wang, Lucy Chai, Xuan Luo, Michael Niemeyer et al.NeurIPS 2025 · 5 citations
- iSplat: Iterative Learning for Fine-Grained Gaussian SplattingHaifeng Wu, Wei Long, Shuhang Gu, Lixin Duan et al.CVPR 2026
