AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation
Luoxi Jing, Dianxi Shi, YuShe Cao, Yuanze Wang, Junze Zhang, Yuning Cui, Mengzhu Wang
Abstract
Monocular depth estimation is critical for applications like autonomous driving and robotics. The complementary properties of event and image modality motivate the fusion-based methods for robust depth estimation. However, existing fusion methods rely on convolutional or attention-based architectures, which either struggle with global dependencies or incur high computational cost, limiting their suitability for long-sequence modeling in depth tasks. Besides, effective image-event fusion remains a key challenge, as most existing methods directly fuse features without addressing the domain gap and differences in representational characteristics between raw events and images, leading to semantic bias and degraded performance. In this work, we propose AIMDepth, an Asymmetric Image-Event Mamba framework for monocular depth estimation, built entirely on state space models to ensure linear computational complexity and accurate prediction. To address input-domain misalignment, we introduce a Spectral Cross-modal Prior Guidance (SCPG) module that performs bidirectional prior injection at the input level. To mitigate representational imbalance between sparse events and dense images, we design an Asymmetric Modal-aware Encoder (AME) that allocates separate encoding paths for each modality and facilitates feature-level alignment tailored to their distinct information densities. To further enhance fusion, we develop a Modality-interactive Local Refinement (ModiLocal) module that enables hierarchical interaction and fine-grained alignment through SSM-based modeling. Extensive experiments on public datasets demonstrate that AIMDepth achieves state-of-the-art performance and strong robustness in complex environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e394db7-ea99-4a32-910a-2f1494d496d6Builds on12
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Learning an Event Sequence Embedding for Dense Event-Based Deep StereoStepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler et al.ICCV 2019 · 122 citations
Related papers
- Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image DomainHanyue Lou, Jinxiu (Sherry) Liang, Minggui Teng, Bin Fan et al.NeurIPS 2024 · 13 citations
- HAFUNet: A Hierarchical Attention Fusion Network for Monocular Depth Estimation Integrating Event and Frame DataSiyuan Zhang, Xiaoping Wang, Jiang Li, Weibin Feng et al.ACM MM 2025
- Distil-E2D: Distilling Image-to-Depth Priors for Event-Based Monocular Depth EstimationJie Long Lee, Gim Hee LeeNeurIPS 2025 · 3 citations
- MambaSeg: Harnessing Mamba for Accurate and Efficient Image-Event Semantic SegmentationFuqiang Gu, Yuanke Li, Xianlei Long, Kangping Ji et al.AAAI 2026 · 1 citation
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu et al.ICCV 2025 · 3 citations
