Long-Tail Internet Photo Reconstruction
Yuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor, Noah Snavely, Ruojin Cai
摘要
Pretrained 𝜋 3 Ours Pretrained 𝜋 3 Ours Scene (sorted by #images) Registered images Total images #images per Scene Long Tail Duomo (Cagliari) -Crypt Calvaire de Plougonven Figure 1. Long-tail Internet photo reconstruction. Internet photo collections follow a long-tailed distribution. In the top plot, the x-axis represents scene index (sorted by image count) and the y-axis shows images per scene (scenes are drawn from MegaScenes [36], a dataset of Internet photo collections). The light blue curve plots the total number of Internet photos per scene, while the steel blue curve shows the size of the subset of photos that were successfully registered using SfM. The head of this distribution of photo collections represents well-photographed scenes; here, there are 6,985 scenes with >50 registered images. However, most photo collections are in the long tail of this distribution; here, 418,056 scenes with fewer than 50 registered photos. State-of-the-art methods often fail on scenes in this tail. In the lower half of the figure, we show two examples from the long tail, along with representative input images and the corresponding reconstructions.
On Calvaire de Plougonven, COLMAP doesn't register any image; on both Duomo (Cagliari)-Crypt and Calvaire de Plougonven, recent feed-forward reconstruction models like π 3 [44] produce poor results. We propose MegaDepth-X dataset and a strategy for mimicking long-tail camera distributions, on which fine-tuned models like π 3 exhibit better reconstruction robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- LightGlue: Local Feature Matching at Light SpeedPhilipp Lindenberger, Paul-Edouard Sarlin, Marc PollefeysICCV 2023 · 被引用 936 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
- DISK: Learning local features with policy gradientMichal J. Tyszkiewicz, Pascal Fua, Eduard TrullsNeurIPS 2020 · 被引用 652 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- Neural RGB-D Surface ReconstructionDejan Azinovic, Ricardo Martin-Brualla, Dan B. Goldman, Matthias Nießner 等CVPR 2022 · 被引用 272 次
相关 Paper
- Sparse-View Localization via Online Neural 3D RegressionLudvig Dillén, Magnus Oskarsson, Viktor LarssonCVPR 2026
- Extreme Rotation Estimation in the WildHana Bezalel, Dotan Ankri, Ruojin Cai, Hadar Averbuch-ElorCVPR 2025
- Geometry of Long-Tailed Representation Learning: Rebalancing Features for Skewed DistributionsLingjie Yi, Jiachen Yao, Weimin Lyu, Haibin Ling 等ICLR 2025
- ULTRA-360: Unconstrained Dataset for Large-scale Temporal 3D Reconstruction across Altitudes and Omnidirectional ViewsXijun Liu, Zhaoliang Zhang, Yuxiang Guo, Yifan Zhou 等ICLR 2026
- Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal VisionXiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah SnavelyICCV 2021 · 被引用 26 次
