Mask Propagation for Efficient Video Semantic Segmentation
Yuetian Weng, Mingfei Han, Haoyu He, Mingjie Li, Lina Yao, Xiaojun Chang, Bohan Zhuang
Abstract
Video Semantic Segmentation (VSS) involves assigning a semantic label to each pixel in a video sequence. Prior work in this field has demonstrated promising results by extending image semantic segmentation models to exploit temporal relationships across video frames; however, these approaches often incur significant computational costs. In this paper, we propose an efficient mask propagation framework for VSS, called MPVSS. Our approach first employs a strong query-based image segmentor on sparse key frames to generate accurate binary masks and class predictions. We then design a flow estimation module utilizing the learned queries to generate a set of segment-aware flow maps, each associated with a mask prediction from the key frame. Finally, the mask-flow pairs are warped to serve as the mask predictions for the non-key frames. By reusing predictions from key frames, we circumvent the need to process a large volume of video frames individually with resource-intensive segmentors, alleviating temporal redundancy and significantly reducing computational costs. Extensive experiments on VSPW and Cityscapes demonstrate that our mask propagation framework achieves SOTA accuracy and efficiency trade-offs. For instance, our best model with Swin-L backbone outperforms the SOTA MRCFA using MiT-B5 by 4.0% mIoU, requiring only 26% FLOPs on the VSPW dataset. Moreover, our framework reduces up to 4x FLOPs compared to the per-frame Mask2Former baseline with only up to 2% mIoU degradation on the Cityscapes validation set. Code is available at https://github.com/ziplab/MPVSS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 594eec2b-299a-416e-a08c-4167b583891bCited by top-tier papers7
- VidEoMT: Your ViT is Secretly Also a Video Segmentation ModelNarges Norouzi, Idil Esen Zulfikar, Niccolò Cavagnero, Tommie Kerssies et al.CVPR 2026 · 8 citations
- End-to-End Video Semantic Segmentation in Adverse Weather using Fusion Blocks and Temporal-Spatial Teacher-Student LearningXin Yang, Wending Yan, Michael Bi Mi, Yuan Yuan et al.NeurIPS 2024 · 6 citations
- Bootstrapping Video Semantic Segmentation Model via Distillation-assisted Test-Time AdaptationJihun Kim, Hoyong Kwon, Hyeokjun Kweon, Kuk-Jin YoonCVPR 2026 · 3 citations
- Dual-Temporal Exemplar Representation Network for Video Semantic SegmentationXiaolong Xu, Lei Zhang, Jiayi Li, Lituan Wang et al.ICCV 2025 · 3 citations
- RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic SegmentationKai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan ZhouCVPR 2026 · 1 citation
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosYubin Hu, Yuze He, Yanghao Li, Jisheng Li et al.CVPR 2023
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi et al.AAAI 2020 · 80 citations
- Accelerating Video Object Segmentation with Compressed VideoKai Xu, Angela YaoCVPR 2022 · 24 citations
- Video Semantic Segmentation via Sparse Temporal TransformerJiangtong Li, Wentao Wang, Junjie Chen, Li Niu et al.ACM MM 2021 · 47 citations
- QueryProp: Object Query Propagation for High-Performance Video Object DetectionFei He, Naiyu Gao, Jian Jia, Xin Zhao et al.AAAI 2022 · 35 citations
