Deep Permutation Equivariant Structure from Motion
Dror Moran, Hodaya Koslowsky, Yoni Kasten, Haggai Maron, Meirav Galun, Ronen Basri
Abstract
Existing deep methods produce highly accurate 3D reconstructions in stereo and multiview stereo settings, i.e., when cameras are both internally and externally calibrated. Nevertheless, the challenge of simultaneous recovery of camera poses and 3D scene structure in multiview settings with deep networks is still outstanding. Inspired by projective factorization for Structure from Motion (SFM) and by deep matrix completion techniques, we propose a neural network architecture that, given a set of point tracks in multiple images of a static scene, recovers both the camera parameters and a (sparse) scene structure by minimizing an unsupervised reprojection loss. Our network architecture is designed to respect the structure of the problem: the sought output is equivariant to permutations of both cameras and scene points. Notably, our method does not require initialization of camera parameters or 3D point locations. We test our architecture in two setups: (1) single scene reconstruction and (2) learning from multiple scenes. Our experiments, conducted on a variety of datasets in both internally calibrated and uncalibrated settings, indicate that our method accurately recovers pose and structure, on par with classical state of the art methods. Additionally, we show that a pre-trained network can be used to reconstruct novel scenes using inexpensive fine-tuning with no loss of accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba328d38-348c-4683-9d08-57a2808b1effCited by top-tier papers8
- Fast Encoder-Based 3D from Casual Videos via Point Track ProcessingYoni Kasten, Wuyue Lu, Haggai MaronNeurIPS 2024 · 16 citations
- OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesTao Xie, Kun Dai, Siyi Lu, Ke Wang et al.ICCV 2023 · 16 citations
- COFS: COntrollable Furniture layout SynthesisWamiq Reyaz Para, Paul Guerrero, Niloy J. Mitra, Peter WonkaSIGGRAPH 2023 · 15 citations
- AETHER: Geometric-Aware Unified World ModelingHaoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang et al.ICCV 2025 · 9 citations
- Consensus Learning with Deep Sets for Essential Matrix EstimationDror Moran, Yuval Margalit, Guy Trostianetsky, Fadi Khatib et al.NeurIPS 2024 · 4 citations
Builds on6
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
- Hierarchical Neural Architecture Search for Deep Stereo MatchingXuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai et al.NeurIPS 2020 · 436 citations
- On Learning Sets of Symmetric ElementsHaggai Maron, Or Litany, Gal Chechik, Ethan FetayaICML 2020 · 148 citations
- Algebraic Characterization of Essential Matrices and Their Averaging in Multiview SettingsYoni Kasten, Amnon Geifman, Meirav Galun, Ronen BasriICCV 2019 · 35 citations
- Synchronizing Probability Measures on Rotations via Optimal TransportTolga Birdal, Michael Arbel, Umut Simsekli, Leonidas J. GuibasCVPR 2020
Related papers
- RESfM: Robust Deep Equivariant Structure from MotionFadi Khatib, Yoni Kasten, Dror Moran, Meirav Galun et al.ICLR 2025
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 72 citations
- VGGSfM: Visual Geometry Grounded Deep Structure from MotionJianyuan Wang, Nikita Karaev, Christian Rupprecht, David NovotnýCVPR 2024 · 48 citations
- Deep Unsupervised 3D SfM Face Reconstruction Based on Massive Landmark Bundle AdjustmentYuxing Wang, Yawen Lu, Zhihua Xie, Guoyu LuACM MM 2021 · 15 citations
- Unsupervised 3D Reconstruction NetworksGeonho Cha, Minsik Lee, Songhwai OhICCV 2019 · 14 citations
