An Embedding-Unleashing Video Polyp Segmentation Framework via Region Linking and Scale Alignment
Zhixue Fang, Xinrong Guo, Jingyin Lin, Huisi Wu, Jing Qin
Abstract
Automatic polyp segmentation from colonoscopy videos is a critical task for the development of computer-aided screening and diagnosis systems. However, accurate and real-time video polyp segmentation (VPS) is a very challenging task due to low contrast between background and polyps and frame-to-frame dramatic variations in colonoscopy videos. We propose a novel embedding-unleashing framework consisting of a proposal-generative network (PGN) and an appearance-embedding network (AEN) to comprehensively address these challenges. Our framework, for the first time, models VPS as an appearance-level semantic embedding process to facilitate generate more global information to counteract background disturbances and dramatic variations. Specifically, PGN is a video segmentation network to obtain segmentation mask proposals, while AEN is a network we specially designed to produce appearance-level embedding semantics for PGN, thereby unleashing the capability of PGN in VPS. Our AEN consists of a cross-scale region linking (CRL) module and a cross-wise scale alignment (CSA) module. The former screens reliable background information against background disturbances by constructing linking of region semantics, while the latter performs the scale alignment to resist dramatic variations by modeling the center-perceived motion dependence with a cross-wise manner. We further introduce a parameter-free semantic interaction to embed the semantics of AEN into PGN to obtain the segmentation results. Extensive experiments on CVC-612 and SUN-SEG demonstrate that our approach achieves better performance than other state-of-the-art methods. Codes are available at https://github.com/zhixue-fang/EUVPS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4773cef-50e6-4acf-b709-7b99e6b8de2dBuilds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Twins: Revisiting the Design of Spatial Attention in Vision TransformersXiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang et al.NeurIPS 2021 · 1,388 citations
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
- VITA: Video Instance Segmentation via Object Token AssociationMiran Heo, Sukjun Hwang, Seoung Wug Oh, Joon-Young Lee et al.NeurIPS 2022 · 146 citations
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai et al.NeurIPS 2021 · 92 citations
Related papers
- WavePolyp: Video Polyp Segmentation via Hierarchical Wavelet-Based Feature Aggregation and Inter-Frame Divergence PerceptionYuhua Zhang, Guilian Chen, Yuanqin He, Huisi Wu et al.ICLR 2026
- HFSTI-Net: Hierarchical Frequency-spatial-temporal Interactions for Video Polyp SegmentationYuanqin He, Guilian Chen, Yuhua Zhang, Huisi Wu et al.ICLR 2026
- VPSentry: Semi-supervised Video Polyp Segmentation via Sentry-guided Long-term Prototype Fusion with Correlation Dynamic PropagationGuilian Chen, Xiaoling Luo, Huisi Wu, Jing QinAAAI 2026
- STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation LearningGuilian Chen, Huisi Wu, Jing QinICCV 2025 · 4 citations
- Precise Yet Efficient Semantic Calibration and Refinement in ConvNets for Real-time Polyp Segmentation from Colonoscopy VideosHuisi Wu, Jiafu Zhong, Wei Wang, Zhenkun Wen et al.AAAI 2021 · 73 citations
