Learning Structure-From-Motion with Graph Attention Networks
Lucas Brynte, José Pedro Iglesias, Carl Olsson, Fredrik Kahl
Abstract
In this paper we tackle the problem of learning Structurefrom-Motion (SfM) through the use of graph attention networks. SfM is a classic computer vision problem that is solved though iterative minimization of reprojection errors, referred to as Bundle Adjustment (BA), starting from a good initialization. In order to obtain a good enough initialization to BA, conventional methods rely on a sequence of sub-problems (such as pairwise pose estimation, pose averaging or triangulation) which provide an initial solution that can then be refined using BA. In this work we replace these sub-problems by learning a model that takes as input the 2D keypoints detected across multiple views, and outputs the corresponding camera poses and 3D keypoint coordinates. Our model takes advantage of graph neural networks to learn SfM-specific primitives, and we show that it can be used for fast inference of the reconstruction for new and unseen sequences. The experimental results show that the proposed model outperforms competing learning-based methods, and challenges COLMAP while having lower runtime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbf4a9a1-d10c-4b71-b70b-27ec489a1c21Cited by top-tier papers6
- VGGSfM: Visual Geometry Grounded Deep Structure from MotionJianyuan Wang, Nikita Karaev, Christian Rupprecht, David NovotnýCVPR 2024 · 48 citations
- Fast Encoder-Based 3D from Casual Videos via Point Track ProcessingYoni Kasten, Wuyue Lu, Haggai MaronNeurIPS 2024 · 16 citations
- MuM: Multi-View Masked Image Modeling for 3D VisionDavid Nordström, Johan Edstedt, Fredrik Kahl, Georg BökmanCVPR 2026 · 6 citations
- Global-Aware Edge Prioritization for Pose Graph InitializationTong Wei, Giorgos Tolias, Jiri Matas, Daniel BarathCVPR 2026 · 1 citation
- Uncalibrated Structure from Motion on a SphereJonathan Ventura, Viktor Larsson, Fredrik KahlICCV 2025 · 1 citation
Builds on7
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view StereoYi Wei, Shaohui Liu, Yongming Rao, Wang Zhao et al.ICCV 2021 · 286 citations
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- Algebraic Characterization of Essential Matrices and Their Averaging in Multiview SettingsYoni Kasten, Amnon Geifman, Meirav Galun, Ronen BasriICCV 2019 · 35 citations
- Deep Permutation Equivariant Structure from MotionDror Moran, Hodaya Koslowsky, Yoni Kasten, Haggai Maron et al.ICCV 2021 · 21 citations
Related papers
- Learning to Bundle-adjust: A Graph Network Approach to Faster Optimization of Bundle Adjustment for Vehicular SLAMTetsuya Tanaka, Yukihiro Sasagawa, Takayuki OkataniICCV 2021 · 9 citations
- Light3R-SfM: Towards Feed-forward Structure-from-MotionSven Elflein, Qunjie Zhou, Laura Leal-TaixéCVPR 2025
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- Level-S2fM: Structure from Motion on Neural Level Set of Implicit SurfacesYuxi Xiao, Nan Xue, Tianfu Wu, Gui-Song XiaCVPR 2023
