Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
Suchisrit Gangopadhyay, Jung Hee Kim, Xien Chen, Patrick Rim, Hyoungseob Park, Alex Wong
摘要
We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite being trained on tens of millions of images, FMDEs are susceptible to the covariate shift introduced by changes in camera calibration (intrinsic, distortion) parameters, leading to erroneous depth estimates. Our method aligns the distribution of latent embeddings encoding fisheye images to those of perspective images, enabling the reuse of FMDEs for fisheye cameras without retraining or finetuning. To this end, we introduce a set of Calibration Tokens as a light-weight adaptation mechanism that modulates the latent embeddings for alignment. By exploiting the already expressive latent space of FMDEs, we posit that modulating their embeddings avoids the negative impact of artifacts and loss introduced in conventional recalibration or map projection to a canonical reference frame in the image space. Our method is self-supervised and does not require fisheye images but leverages publicly available large-scale perspective image datasets. This is done by recalibrating perspective images to fisheye images, and enforcing consistency between their estimates during training. We evaluate our approach with several FMDEs, on both indoors and outdoors, where we consistently improve over state-of-the-art methods using a single set of tokens for both. Code available at: https://github.com/JungHeeKim29/calibration-token; https://github.com/Suchisrit/CalibrationTokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- UniDAC: Universal Metric Depth Estimation for Any CameraGirish Chandar Ganesan, Yuliang Guo, Liu Ren, Xiaoming LiuCVPR 2026 · 被引用 8 次
- Radar-Guided Polynomial Fitting for Metric Depth EstimationPatrick Rim, Hyoungseob Park, Vadim Ezhov, Jeffrey Moon 等CVPR 2026 · 被引用 7 次
- ETA: Energy-Based Test-Time Adaptation for Depth CompletionYounjoon Chung, Hyoungseob Park, Patrick Rim, Xiaoran Zhang 等ICCV 2025 · 被引用 1 次
- ORCaS: Unsupervised Depth Completion via Occluded Region Completion as SupervisionHyoungseob Park, Runjian Chen, Patrick Rim, Dong Lao 等ICLR 2026
- Entropy-Monitored Kernelized Token Distillation for Audio-Visual CompressionHyoungseob Park, Lipeng Ke, Pritish Mohapatra, Huajun Ying 等ICLR 2026
它引用的顶会 Paper47
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
相关 Paper
- SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation LearningHao Feng, Wendi Wang, Jiajun Deng, Wengang Zhou 等ICCV 2023 · 被引用 28 次
- Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric VisionTianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar 等ICCV 2025 · 被引用 1 次
- Estimating Egocentric 3D Human Pose in Global SpaceJian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar 等ICCV 2021 · 被引用 78 次
- Structure-from-Motion with a Non-Parametric Camera ModelYihan Wang, Linfei Pan, Marc Pollefeys, Viktor LarssonCVPR 2025
- Deep Single Image Camera Calibration by Heatmap Regression to Recover Fisheye Images Under Manhattan World AssumptionNobuhiko Wakai, Satoshi Sato, Yasunori Ishii, Takayoshi YamashitaCVPR 2024 · 被引用 4 次
