LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation models
Ziqi Lu, Heng Yang, Danfei Xu, Boyi Li, Boris Ivanovic, Marco Pavone, Yue Wang
摘要
Emerging 3D geometric foundation models, such as DUSt3R (Wang et al., 2024) , offer a promising approach for in-the-wild 3D vision tasks. However, due to the high-dimensional nature of the problem space and scarcity of high-quality 3D data, these pre-trained models still struggle to generalize to many challenging circumstances, such as limited view overlap or low lighting. To address this, we propose LoRA3D, an efficient self-calibration pipeline to specialize the pre-trained models to target scenes using their own multi-view predictions. Taking sparse RGB images as input, we leverage robust optimization techniques to refine multiview predictions and align them into a global coordinate frame. In particular, we incorporate prediction confidence into the geometric optimization process, automatically re-weighting the confidence to better reflect point estimation accuracy. We use the calibrated confidence to generate high-quality pseudo labels for the calibrating views and use low-rank adaptation (LoRA) to fine-tune the models on the pseudo-labeled data. Our method does not require any external priors or manual labels. It completes the self-calibration process on a single standard GPU within just 5 minutes. Each low-rank adapter requires only 18MB of storage. We evaluated our method on more than 160 scenes from the Replica, TUM and Waymo Open datasets, achieving up to 88% performance improvement on 3D reconstruction, multi-view pose estimation and novel-view rendering. For more details, please visit our project page. Published as a conference paper at ICLR 2025 Recently, many 3D geometric foundation models have emerged as a potential solution for in-thewild 3D computer vision tasks such as 3D reconstruction, camera pose estimation and novel view rendering (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 被引用 17 次
- TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable TokensJiawei Ren, Michal J. Tyszkiewicz, Jiahui Huang, Zan GojcicCVPR 2026 · 被引用 13 次
- G4Splat: Geometry-Guided Gaussian Splatting with Generative PriorJunfeng Ni, Yixin Chen, Zhifei Yang, Yu Liu 等ICLR 2026 · 被引用 10 次
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge DistillationWeining Ren, Hongjun Wang, Xiao Tan, Kai HanNeurIPS 2025 · 被引用 5 次
- LONG3R: Long Sequence Streaming 3D ReconstructionZhuoguang Chen, Minghui Qin, Tianyuan Yuan, Zhe Liu 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- iMAP: Implicit Mapping and Positioning in Real-TimeEdgar Sucar, Shikun Liu, Joseph Ortiz, Andrew J. DavisonICCV 2021 · 被引用 834 次
相关 Paper
- Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionWenyu Li, Sidun Liu, Peng Qiao, Yong DouACM MM 2025 · 被引用 3 次
- Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation ModelsYu Jiang, Hanwen Jiang, Ahmed Abdelkader, Wen-Sheng Chu 等CVPR 2026
- Dr. RAW: Towards General High-Level Vision from RAW with Efficient Task ConditioningWenjun Huang, Ziteng Cui, Yinqiang Zheng, Yirui He 等NeurIPS 2025 · 被引用 5 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature AlignmentYouming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal 等CVPR 2026 · 被引用 4 次
