Depth AnyEvent: A Cross-Modal Distillation Paradigm for Event-Based Monocular Depth Estimation
Luca Bartolomei, Enrico Mannocci, Fabio Tosi, Matteo Poggi, Stefano Mattoccia
Abstract
Event cameras capture sparse, high-temporal-resolution visual information, making them particularly suitable for challenging environments with high-speed motion and strongly varying lighting conditions. However, the lack of large datasets with dense ground-truth depth annotations hinders learning-based monocular depth estimation from event data. To address this limitation, we propose a cross-modal distillation paradigm to generate dense proxy labels leveraging a Vision Foundation Model (VFM). Our strategy requires an event stream spatially aligned with RGB frames, a simple setup even available off-the-shelf, and exploits the robustness of large-scale VFMs. Additionally, we propose to adapt VFMs, either a vanilla one like Depth Anything v2 (DAv2), or deriving from it a novel recurrent architecture to infer depth from monocular event cameras. We evaluate our approach with synthetic and real-world datasets, demonstrating that i) our cross-modal paradigm achieves competitive performance compared to fully supervised methods without requiring expensive depth annotations, and ii) our VFM-based models achieve state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ede6619-83ef-45b6-ae3e-54606242d346Cited by top-tier papers4
- Scaling Dense Event-Stream Pretraining from Visual Foundation ModelsZhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu et al.CVPR 2026 · 2 citations
- Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric StereoNinghui Xu, Fabio Tosi, Lihui Wang, Jiawei Han et al.CVPR 2026
- Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth EstimationDaikun Liu, Teng Wang, Changyin SunCVPR 2026
- EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active SensorsLuca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia et al.CVPR 2026
Builds on8
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
- DDP: Diffusion Model for Dense Visual PredictionYuanfeng Ji, Zhe Chen, Enze Xie, Lanqing Hong et al.ICCV 2023 · 223 citations
Related papers
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu et al.ICCV 2025 · 3 citations
- Distil-E2D: Distilling Image-to-Depth Priors for Event-Based Monocular Depth EstimationJie Long Lee, Gim Hee LeeNeurIPS 2025 · 3 citations
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu et al.ICCV 2025 · 1 citation
- DERD-Net: Learning Depth from Event-based Ray DensitiesDiego de Oliveira Hitzges, Suman Ghosh, Guillermo GallegoNeurIPS 2025 · 6 citations
- EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge DistillationLin Wang, Yujeong Chae, Sung-Hoon Yoon, Tae-Kyun Kim et al.CVPR 2021
