Unified Mask Embedding and Correspondence Learning for Self-Supervised Video Segmentation
Liulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li, Yi Yang
Abstract
The objective of this paper is self-supervised learning of video object segmentation. We develop a unified framework which simultaneously models cross-frame dense correspondence for locally discriminative feature learning and embeds object-level context for target-mask decoding. As a result, it is able to directly learn to perform mask-guided sequential segmentation from unlabeled videos, in contrast to previous efforts usually relying on an oblique solution -cheaply "copying" labels according to pixel-wise correlations. Concretely, our algorithm alternates between i) clustering video pixels for creating pseudo segmentation labels ex nihilo; and ii) utilizing the pseudo labels to learn mask encoding and decoding for VOS. Unsupervised correspondence learning is further incorporated into this self-taught, mask embedding scheme, so as to ensure the generic nature of the learnt representation and avoid cluster degeneracy. Our algorithm sets state-of-the-arts on two standard benchmarks (i.e., DAVIS 17 and YouTube-VOS), narrowing the gap between self-and fully-supervised VOS, in terms of both performance and network architecture design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a6cce3c-0717-4cb8-9de7-fa3433624f31Cited by top-tier papers8
- CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationKexin Li, Zongxin Yang, Lei Chen, Yi Yang et al.ACM MM 2023 · 58 citations
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 52 citations
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao et al.ICCV 2023 · 42 citations
- VideoMAC: Video Masked Autoencoders Meet ConvNetsGensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu et al.CVPR 2024 · 14 citations
- O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric ViewsLorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus et al.ICCV 2025 · 2 citations
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
Related papers
- Boosting Video Object Segmentation via Space-Time Correspondence LearningYurong Zhang, Liulei Li, Wenguan Wang, Rong Xie et al.CVPR 2023
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 38 citations
- Video Object Segmentation Using Global and Instance Embedding LearningWenbin Ge, Xiankai Lu, Jianbing ShenCVPR 2021
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 23 citations
- In-N-Out Generative Learning for Dense Unsupervised Video SegmentationXiao Pan, Peike Li, Zongxin Yang, Huiling Zhou et al.ACM MM 2022 · 8 citations
