AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation
Luoxi Jing, Dianxi Shi, YuShe Cao, Yuanze Wang, Junze Zhang, Yuning Cui, Mengzhu Wang
摘要
Monocular depth estimation is critical for applications like autonomous driving and robotics. The complementary properties of event and image modality motivate the fusion-based methods for robust depth estimation. However, existing fusion methods rely on convolutional or attention-based architectures, which either struggle with global dependencies or incur high computational cost, limiting their suitability for long-sequence modeling in depth tasks. Besides, effective image-event fusion remains a key challenge, as most existing methods directly fuse features without addressing the domain gap and differences in representational characteristics between raw events and images, leading to semantic bias and degraded performance. In this work, we propose AIMDepth, an Asymmetric Image-Event Mamba framework for monocular depth estimation, built entirely on state space models to ensure linear computational complexity and accurate prediction. To address input-domain misalignment, we introduce a Spectral Cross-modal Prior Guidance (SCPG) module that performs bidirectional prior injection at the input level. To mitigate representational imbalance between sparse events and dense images, we design an Asymmetric Modal-aware Encoder (AME) that allocates separate encoding paths for each modality and facilitates feature-level alignment tailored to their distinct information densities. To further enhance fusion, we develop a Modality-interactive Local Refinement (ModiLocal) module that enables hierarchical interaction and fine-grained alignment through SSM-based modeling. Extensive experiments on public datasets demonstrate that AIMDepth achieves state-of-the-art performance and strong robustness in complex environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Learning an Event Sequence Embedding for Dense Event-Based Deep StereoStepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler 等ICCV 2019 · 被引用 122 次
相关 Paper
- Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image DomainHanyue Lou, Jinxiu (Sherry) Liang, Minggui Teng, Bin Fan 等NeurIPS 2024 · 被引用 13 次
- HAFUNet: A Hierarchical Attention Fusion Network for Monocular Depth Estimation Integrating Event and Frame DataSiyuan Zhang, Xiaoping Wang, Jiang Li, Weibin Feng 等ACM MM 2025
- Distil-E2D: Distilling Image-to-Depth Priors for Event-Based Monocular Depth EstimationJie Long Lee, Gim Hee LeeNeurIPS 2025 · 被引用 3 次
- MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic SegmentationFuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji 等AAAI 2026 · 被引用 1 次
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu 等ICCV 2025 · 被引用 3 次
