STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation Learning
Guilian Chen, Huisi Wu, Jing Qin
摘要
Automated segmentation of polyps from colonoscopy videos is of great clinical significance as it can assist clinicians in making accurate diagnoses and precise interventions. However, video polyp segmentation (VPS) is challenging due to ambiguous polyp boundaries, as well as variations in polyp scale, contrast, and position across consecutive frames. Moreover, to meet clinical requirements, the inference must operate in real-time to enable intraoperative tracking and guidance. In this paper, we propose a novel and efficient segmentation network, STDDNet, which integrates a spatial-aligned temporal modeling strategy and a discriminative dynamic representation learning mechanism, to comprehensively address these challenges by harnessing the advantages of Mamba. Specifically, a spatialaligned temporal dependency propagation (STDP) module is developed to model temporal consistency from the consecutive frames based on a bidirectional scanning Mamba block. Furthermore, we design a discriminative dynamic feature extraction (DDFE) module to explore frame-wise dynamic information from the structural feature generated by the Mamba block. Such dynamic features can effectively deal with the variations across colonoscopy frames, providing more details for refined segmentation. We extensively evaluate STDDNet on two benchmark datasets, SUN-SEG and CVC-ClinicDB, demonstrating superior segmentation performance compared to state-of-the-art methods while maintaining real-time inference. Codes are available at https://github.com/C-GLGLGL/STDDNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical VideosJialun Pei, Zhangjun Zhou, Diandian Guo, Zhixi Li 等CVPR 2026 · 被引用 6 次
- Clinically-Grounded Counterfactual Reasoning for Medical Video DiagnosisJianzhe Gao, Churan Wang, Weiyi Zhang, Jianghua Li 等CVPR 2026
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Precise Yet Efficient Semantic Calibration and Refinement in ConvNets for Real-time Polyp Segmentation from Colonoscopy VideosHuisi Wu, Jiafu Zhong, Wei Wang, Zhenkun Wen 等AAAI 2021 · 被引用 73 次
- MambaOut: Do We Really Need Mamba for Vision?Weihao Yu, Xinchao WangCVPR 2025
相关 Paper
- VPSentry: Semi-supervised Video Polyp Segmentation via Sentry-guided Long-term Prototype Fusion with Correlation Dynamic PropagationGuilian Chen, Xiaoling Luo, Huisi Wu, Jing QinAAAI 2026
- HFSTI-Net: Hierarchical Frequency-spatial-temporal Interactions for Video Polyp SegmentationYuanqin He, Guilian Chen, Yuhua Zhang, Huisi Wu 等ICLR 2026
- WavePolyp: Video Polyp Segmentation via Hierarchical Wavelet-Based Feature Aggregation and Inter-Frame Divergence PerceptionYuhua Zhang, Guilian Chen, Yuanqin He, Huisi Wu 等ICLR 2026
- An Embedding-Unleashing Video Polyp Segmentation Framework via Region Linking and Scale AlignmentZhixue Fang, Xinrong Guo, Jingyin Lin, Huisi Wu 等AAAI 2024 · 被引用 8 次
- When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object DetectionQiang Qi, Xiao Wang, Zongyuan Du, Yu ZhangCVPR 2026
