Dense Unsupervised Learning for Video Segmentation
Nikita Araslanov, Simone Schaub-Meyer, Stefan Roth
Abstract
We present a novel approach to unsupervised learning for video object segmentation (VOS). Unlike previous work, our formulation allows to learn dense feature representations directly in a fully convolutional regime. We rely on uniform grid sampling to extract a set of anchors and train our model to disambiguate between them on both inter-and intra-video levels. However, a naive scheme to train such a model results in a degenerate solution. We propose to prevent this with a simple regularisation scheme, accommodating the equivariance property of the segmentation task to similarity transformations. Our training objective admits efficient implementation and exhibits fast training convergence. On established VOS benchmarks, our approach exceeds the segmentation accuracy of previous work despite using significantly less training data and compute power. Code (Apache-2.0 License) available at https://github.com/visinf/dense-ulearn-vos . 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f48145f0-ffa4-40bd-bbb9-d5ddeb3c2767Cited by top-tier papers12
- Time Does Tell: Self-Supervised Time-Tuning of Dense Image RepresentationsMohammadreza Salehi, Efstratios Gavves, Cees G. M. Snoek, Yuki M. AsanoICCV 2023 · 34 citations
- Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsZheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin et al.AAAI 2023 · 22 citations
- In-N-Out Generative Learning for Dense Unsupervised Video SegmentationXiao Pan, Peike Li, Zongxin Yang, Huiling Zhou et al.ACM MM 2022 · 8 citations
- Learning Fine-Grained Features for Pixel-wise Video CorrespondencesRui Li, Shenglong Zhou, Dong LiuICCV 2023 · 7 citations
- From ViT Features to Training-free Video Object Segmentation via Streaming-data Mixture ModelsRoy Uziel, Or Dinari, Oren FreifeldNeurIPS 2023 · 6 citations
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- Learning Video Object Segmentation From Unlabeled VideosXiankai Lu, Wenguan Wang, Jianbing Shen, Yu-Wing Tai et al.CVPR 2020
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- Unsupervised Hierarchical Semantic Segmentation with Multiview Cosegmentation and Clustering TransformersTsung-Wei Ke, Jyh-Jing Hwang, Yunhui Guo, Xudong Wang et al.CVPR 2022 · 34 citations
- VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAEHaonan Yu, Wei XuICLR 2024 · 1 citation
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric VideoYuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li et al.CVPR 2026 · 1 citation
