DERD-Net: Learning Depth from Event-based Ray Densities
Diego de Oliveira Hitzges, Suman Ghosh, Guillermo Gallego
摘要
Event cameras offer a promising avenue for multi-view stereo depth estimation and Simultaneous Localization And Mapping (SLAM) due to their ability to detect blur-free 3D edges at high-speed and over broad illumination conditions. However, traditional deep learning frameworks designed for conventional cameras struggle with the asynchronous, stream-like nature of event data, as their architectures are optimized for discrete, image-like inputs. We propose a scalable, flexible and adaptable framework for pixel-wise depth estimation with event cameras in both monocular and stereo setups. The 3D scene structure is encoded into disparity space images (DSIs), representing spatial densities of rays obtained by back-projecting events into space via known camera poses. Our neural network processes local subregions of the DSIs combining 3D convolutions and a recurrent structure to recognize valuable patterns for depth prediction. Local processing enables fast inference with full parallelization and ensures constant ultra-low model complexity and memory costs, regardless of camera resolution. Experiments on standard benchmarks (MVSEC and DSEC datasets) demonstrate unprecedented effectiveness: (i) using purely monocular data, our method achieves comparable results to existing stereo methods; (ii) when applied to stereo data, it strongly outperforms all state-of-the-art (SOTA) approaches, reducing the mean absolute error by at least 42%; (iii) our method also allows for increases in depth completeness by more than 3-fold while still yielding a reduction in median absolute error of at least 30%. Given its remarkable performance and effective processing of event-data, our framework holds strong potential to become a standard approach for using deep learning for event-based depth estimation and SLAM. Project page: https://github.com/tub-rip/DERD-Net
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Event-Driven Dynamic Scene Depth CompletionZhiqiang Yan, Jianhao Jiao, Zhengxue Wang, Gim Hee LeeNeurIPS 2025 · 被引用 12 次
- EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active SensorsLuca Bartolomei, Fabio Tosi, Matteo Poggi, Stefano Mattoccia 等CVPR 2026
它引用的顶会 Paper7
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy 等NeurIPS 2022 · 被引用 183 次
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
- Learning an Event Sequence Embedding for Dense Event-Based Deep StereoStepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler 等ICCV 2019 · 被引用 122 次
- Deep Event Stereo Leveraged by Event-to-Image TranslationSoikat Hasan Ahmed, Hae Woong Jang, S. M. Nadim Uddin, Yong Ju JungAAAI 2021 · 被引用 41 次
- Discrete time convolution for fast event-based stereoKaixuan Zhang, Kaiwei Che, Jianguo Zhang, Jie Cheng 等CVPR 2022 · 被引用 34 次
相关 Paper
- Event-Image Fusion Stereo Using Cross-Modality Feature PropagationHoonhee Cho, Kuk-Jin YoonAAAI 2022 · 被引用 34 次
- Active Event-based Stereo VisionJianing Li, Yunjian Zhang, Haiqian Han, Xiangyang JiCVPR 2025
- Unsupervised 3d Motion Estimation Using Event CameraHan Han, Wei Zhai, Tiesong Zhao, Bin Li 等CVPR 2026
- Enhanced Event-Based Dense Stereo via Cross-Sensor Knowledge DistillationHaihao Zhang, Yunjian Zhang, Jianing Li, Lin Zhu 等ICCV 2025 · 被引用 1 次
- Unleashing the Temporal Potential of Stereo Event Cameras for Continuous-Time 3D Object DetectionJae-Young Kang, Hoonhee Cho, Kuk-Jin YoonICCV 2025 · 被引用 4 次
