Unified Mask Embedding and Correspondence Learning for Self-Supervised Video Segmentation
Liulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li, Yi Yang
摘要
The objective of this paper is self-supervised learning of video object segmentation. We develop a unified framework which simultaneously models cross-frame dense correspondence for locally discriminative feature learning and embeds object-level context for target-mask decoding. As a result, it is able to directly learn to perform mask-guided sequential segmentation from unlabeled videos, in contrast to previous efforts usually relying on an oblique solution -cheaply "copying" labels according to pixel-wise correlations. Concretely, our algorithm alternates between i) clustering video pixels for creating pseudo segmentation labels ex nihilo; and ii) utilizing the pseudo labels to learn mask encoding and decoding for VOS. Unsupervised correspondence learning is further incorporated into this self-taught, mask embedding scheme, so as to ensure the generic nature of the learnt representation and avoid cluster degeneracy. Our algorithm sets state-of-the-arts on two standard benchmarks (i.e., DAVIS 17 and YouTube-VOS), narrowing the gap between self-and fully-supervised VOS, in terms of both performance and network architecture design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationKexin Li, Zongxin Yang, Lei Chen, Yi Yang 等ACM MM 2023 · 被引用 58 次
- LogicSeg: Parsing Visual Semantics with Neural Logic Learning and ReasoningLiulei Li, Wenguan Wang, Yang YiICCV 2023 · 被引用 52 次
- HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object SegmentationMingfei Han, Yali Wang, Zhihui Li, Lina Yao 等ICCV 2023 · 被引用 42 次
- VideoMAC: Video Masked Autoencoders Meet ConvNetsGensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu 等CVPR 2024 · 被引用 14 次
- O-MaMa: Learning Object Mask Matching Between Egocentric and Exocentric ViewsLorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Pérez-Yus 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
相关 Paper
- Boosting Video Object Segmentation via Space-Time Correspondence LearningYurong Zhang, Liulei Li, Wenguan Wang, Rong Xie 等CVPR 2023
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 被引用 38 次
- Video Object Segmentation Using Global and Instance Embedding LearningWenbin Ge, Xiankai Lu, Jianbing ShenCVPR 2021
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 被引用 23 次
- In-N-Out Generative Learning for Dense Unsupervised Video SegmentationXiao Pan, Peike Li, Zongxin Yang, Huiling Zhou 等ACM MM 2022 · 被引用 8 次
