LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
Hui Liu, Chen Jia, Fan Shi, Xu Cheng, Mengfei Shi, Xia Xie, Shengyong Chen
Abstract
Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 690 citations
- It's Raw! Audio Generation with State-Space ModelsKaran Goel, Albert Gu, Chris Donahue, Christopher RéICML 2022 · 257 citations
- CrackFormer: Transformer Network for Fine-Grained Crack DetectionHuajun Liu, Xiangyu Miao, Christoph Mertz, Chengzhong Xu et al.ICCV 2021 · 195 citations
Related papers
- SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in StructuresHui Liu, Chen Jia, Fan Shi, Xu Cheng et al.CVPR 2025
- LFMamba: Focal Stack-aware State Space Modeling for Light Field Salient Object DetectionXinbo Geng, Fan Shi, Xu Cheng, Chen Jia et al.ACM MM 2025
- CrackSSM: Reviving SSMs for Crack Segmentation via Dynamic ScanningYubin Gu, Boyang Hou, Yuan Meng, Wenting Luo et al.CVPR 2026
- MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba AttentionZilong Zhao, Zhengming Ding, Pei Niu, Wenhao Sun et al.CVPR 2026 · 12 citations
- A Novel State Space Model with Local Enhancement and State Sharing for Image FusionZihan Cao, Xiao Wu, Liang-Jian Deng, Yu ZhongACM MM 2024 · 24 citations
