Simultaneously Short- and Long-Term Temporal Modeling for Semi-Supervised Video Semantic Segmentation
Jiangwei Lao, Weixiang Hong, Xin Guo, Yingying Zhang, Jian Wang, Jingdong Chen, Wei Chu
摘要
In order to tackle video semantic segmentation task at a lower cost, e.g., only one frame annotated per video, lots of efforts have been devoted to investigate the utilization of those unlabeled frames by either assigning pseudo labels or performing feature enhancement. In this work, we propose a novel feature enhancement network to simultaneously model short-and long-term temporal correlation. Compared with existing work that only leverage short-term correspondence, the long-term temporal correlation obtained from distant frames can effectively expand the temporal perception field and provide richer contextual prior. More importantly, modeling adjacent and distant frames together can alleviate the risk of over-fitting, hence produce high-quality feature representation for the distant unlabeled frames in training set and unseen videos in testing set. To this end, we term our method SSLTM, short for Simultaneously Shortand Long-Term Temporal Modeling. In the setting of only one frame annotated per video, SSLTM significantly outperforms the state-of-the-art methods by 2% ∼ 3% mIoU on the challenging VSPW dataset. Furthermore, when working with a pseudo label based method such as MeanTeacher, our final model only exhibits 0.13% mIoU less than the ceiling performance (i.e., all frames are manually annotated).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic SegmentationKai Zhu, Zhenyu Cui, Zehua Zang, Jiahuan ZhouCVPR 2026 · 被引用 1 次
- High Temporal Consistency through Semantic Similarity Propagation in Semi-Supervised Video Semantic Segmentation for Autonomous FlightCédric Vincent, Taehyoung Kim, Henri MeeßCVPR 2025
- Exploiting Temporal State Space Sharing for Video Semantic SegmentationSyed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding 等CVPR 2025
它引用的顶会 Paper11
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- PseudoSeg: Designing Pseudo Labels for Semantic SegmentationYuliang Zou, Zizhao Zhang, Han Zhang, Chun-Liang Li 等ICLR 2021 · 被引用 364 次
相关 Paper
- Every Frame Counts: Joint Learning of Video Segmentation and Optical FlowMingyu Ding, Zhe Wang, Bolei Zhou, Jianping Shi 等AAAI 2020 · 被引用 80 次
- Minimizing Labeled, Maximizing Unlabeled: An Image-Driven Approach for Video Instance SegmentationFangyun Wei, Jinjing Zhao, Kun Yan, Chang XuCVPR 2025
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li 等CVPR 2023
- Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic SegmentationMatthieu Lin, Jenny Sheng, Yubin Hu, Yangguang Li 等AAAI 2024 · 被引用 3 次
- Semi-Supervised Video Semantic Segmentation with Inter-Frame Feature ReconstructionJiafan Zhuang, Zilei Wang, Yuan GaoCVPR 2022 · 被引用 14 次
