ScaleNet: A Shallow Architecture for Scale Estimation
Axel Barroso Laguna, Yurun Tian, Krystian Mikolajczyk
Abstract
In this paper, we address the problem of estimating scale factors between images. We formulate the scale estimation problem as a prediction of a probability distribution over scale factors. We design a new architecture, SealeNet, that exploits dilated convolutions as well as self- and cross-correlation layers to predict the scale between images. We demonstrate that rectifying images with estimated scales leads to significant performance improvements for various tasks and methods. Specifically, we show how ScaleNet can be combined with sparse local features and dense correspondence networks to improve camera pose estimation, 3D reconstruction, or dense geometric matching in different benchmarks and datasets. We provide an extensive evaluation on several tasks, and analyze the computational overhead of SealeNet. The code, evaluation protocols, and trained models are publicly available at https://github.com/axelBarroso/ScaleNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- ETO: Efficient Transformer-based Local Feature Matching by Organizing Multiple Homography HypothesesJunjie Ni, Guofeng Zhang, Guanglin Li, Yijin Li et al.NeurIPS 2024 · 14 citations
- A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image FeaturesAxel Barroso-Laguna, Tommaso Cavallari, Victor Prisacariu, Eric BrachmannICLR 2026 · 3 citations
- Two-View Geometry Scoring Without CorrespondencesAxel Barroso-Laguna, Eric Brachmann, Victor Adrian Prisacariu, Gabriel J. Brostow et al.CVPR 2023
- Matching 2D Images in 3D: Metric Relative Pose from Metric CorrespondencesAxel Barroso-Laguna, Sowmya Munukutla, Victor Adrian Prisacariu, Eric BrachmannCVPR 2024
- Adaptive Assignment for Geometry Aware Local Feature MatchingDihe Huang, Ying Chen, Yong Liu, Jianlin Liu et al.CVPR 2023
Builds on10
- COTR: Correspondence Transformer for Matching Across ImagesWei Jiang, Eduard Trulls, Jan Hosang, Andrea Tagliasacchi et al.ICCV 2021 · 318 citations
- Neural-Guided RANSAC: Learning Where to Sample Model HypothesesEric Brachmann, Carsten RotherICCV 2019 · 282 citations
- GOCor: Bringing Globally Optimized Correspondence Volumes into Your Neural NetworkPrune Truong, Martin Danelljan, Luc Van Gool, Radu TimofteNeurIPS 2020 · 89 citations
- Beyond Cartesian Representations for Local DescriptorsPatrick Ebel, Eduard Trulls, Kwang Moo Yi, Pascal Fua et al.ICCV 2019 · 83 citations
- ASLFeat: Learning Local Features of Accurate Shape and LocalizationZixin Luo, Lei Zhou, Xuyang Bai, Hongkai Chen et al.CVPR 2020
Related papers
- Learning Camera Localization via Dense Scene MatchingShitao Tang, Chengzhou Tang, Rui Huang, Siyu Zhu et al.CVPR 2021
- Learning Affine Correspondences by Integrating Geometric ConstraintsPengju Sun, Banglei Guan, Zhenbao Yu, Yang Shang et al.CVPR 2025
- ScaleNet - Improve CNNs through Recursively Rescaling ObjectsXingyi Li, Zhongang Qi, Xiaoli Z. Fern, Fuxin LiAAAI 2020 · 1 citation
- FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationChris Rockwell, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park et al.CVPR 2024
- GLU-Net: Global-Local Universal Network for Dense Flow and CorrespondencesPrune Truong, Martin Danelljan, Radu TimofteCVPR 2020
