STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation Learning
Guilian Chen, Huisi Wu, Jing Qin
Abstract
Automated segmentation of polyps from colonoscopy videos is of great clinical significance as it can assist clinicians in making accurate diagnoses and precise interventions. However, video polyp segmentation (VPS) is challenging due to ambiguous polyp boundaries, as well as variations in polyp scale, contrast, and position across consecutive frames. Moreover, to meet clinical requirements, the inference must operate in real-time to enable intraoperative tracking and guidance. In this paper, we propose a novel and efficient segmentation network, STDDNet, which integrates a spatial-aligned temporal modeling strategy and a discriminative dynamic representation learning mechanism, to comprehensively address these challenges by harnessing the advantages of Mamba. Specifically, a spatialaligned temporal dependency propagation (STDP) module is developed to model temporal consistency from the consecutive frames based on a bidirectional scanning Mamba block. Furthermore, we design a discriminative dynamic feature extraction (DDFE) module to explore frame-wise dynamic information from the structural feature generated by the Mamba block. Such dynamic features can effectively deal with the variations across colonoscopy frames, providing more details for refined segmentation. We extensively evaluate STDDNet on two benchmark datasets, SUN-SEG and CVC-ClinicDB, demonstrating superior segmentation performance compared to state-of-the-art methods while maintaining real-time inference. Codes are available at https://github.com/C-GLGLGL/STDDNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f77374a6-69bb-434a-9819-5c0764f8466aCited by top-tier papers2
- Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical VideosJialun Pei, Zhangjun Zhou, Diandian Guo, Zhixi Li et al.CVPR 2026 · 6 citations
- Clinically-Grounded Counterfactual Reasoning for Medical Video DiagnosisJianzhe Gao, Churan Wang, Weiyi Zhang, Jianghua Li et al.CVPR 2026
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Precise Yet Efficient Semantic Calibration and Refinement in ConvNets for Real-time Polyp Segmentation from Colonoscopy VideosHuisi Wu, Jiafu Zhong, Wei Wang, Zhenkun Wen et al.AAAI 2021 · 73 citations
- MambaOut: Do We Really Need Mamba for Vision?Weihao Yu, Xinchao WangCVPR 2025
Related papers
- VPSentry: Semi-supervised Video Polyp Segmentation via Sentry-guided Long-term Prototype Fusion with Correlation Dynamic PropagationGuilian Chen, Xiaoling Luo, Huisi Wu, Jing QinAAAI 2026
- HFSTI-Net: Hierarchical Frequency-spatial-temporal Interactions for Video Polyp SegmentationYuanqin He, Guilian Chen, Yuhua Zhang, Huisi Wu et al.ICLR 2026
- WavePolyp: Video Polyp Segmentation via Hierarchical Wavelet-Based Feature Aggregation and Inter-Frame Divergence PerceptionYuhua Zhang, Guilian Chen, Yuanqin He, Huisi Wu et al.ICLR 2026
- An Embedding-Unleashing Video Polyp Segmentation Framework via Region Linking and Scale AlignmentZhixue Fang, Xinrong Guo, Jingyin Lin, Huisi Wu et al.AAAI 2024 · 8 citations
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
