Semi-Supervised Video Semantic Segmentation with Inter-Frame Feature Reconstruction
Jiafan Zhuang, Zilei Wang, Yuan Gao
Abstract
One major challenge for semantic segmentation in realworld scenarios is only limited pixel-level labels available due to high expense of human labor though a vast volume of video data is provided. Existing semi-supervised methods attempt to exploit unlabeled data in model training, but they just regard video as a set of independent images. To better explore semi-supervised segmentation problem with video data, we formulate a semi-supervised video semantic segmentation task in this paper. For this task, we observe that the overfitting is surprisingly severe between labeled and unlabeled frames within a training video although they are very similar in style and contents. This is called inner-video overfitting, and it would actually lead to inferior performance. To tackle this issue, we propose a novel interframe feature reconstruction (IFR) technique to leverage the ground-truth labels to supervise the model training on unlabeled frames. IFR is essentially to utilize the internal relevance of different frames within a video. During training, IFR would enforce the feature distributions between labeled and unlabeled frames to be narrowed. Consequently, the inner-video overfitting issue can be effectively alleviated. We conduct extensive experiments on Cityscapes and CamVid, and the results demonstrate the superiority of our proposed method to previous state-of-the-art methods. The code is available at https://github.com/jfzhuang/IFR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9444adf-0d45-4287-86be-cb4af07c140bCited by top-tier papers6
- Unleashing Multispectral Video's Potential in Semantic Segmentation: A Semi-supervised Viewpoint and New UAV-View BenchmarkWei Ji, Jingjing Li, Wenbo Li, Yilin Shen et al.NeurIPS 2024 · 8 citations
- Infer from What You Have Seen Before: Temporally-dependent Classifier for Semi-supervised Video SegmentationJiafan Zhuang, Zilei Wang, Yixin Zhang, Zhun FanCVPR 2024 · 4 citations
- From Sharp to Blur: Unsupervised Domain Adaptation for 2D Human Pose Estimation Under Extreme Motion Blur Using Event CamerasYoungho Kim, Hoonhee Cho, Kuk-Jin YoonICCV 2025 · 2 citations
- High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous FlightCédric Vincent, Taehyoung Kim, Henri MeeßCVPR 2025
- Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic SegmentationJiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang et al.CVPR 2023
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi et al.AAAI 2020 · 80 citations
- Siamese Network with Interactive Transformer for Video Object SegmentationMeng Lan, Jing Zhang, Fengxiang He, Lefei ZhangAAAI 2022 · 41 citations
- OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware InterpolationJisoo Jeong, Hong Cai, Risheek Garrepalli, Jamie Menjay Lin et al.CVPR 2024
- Unsupervised Video Object Segmentation with Online Adversarial Self-TuningTiankang Su, Huihui Song, Dong Liu, Bo Liu et al.ICCV 2023 · 19 citations
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 38 citations
