Learning Fine-Grained Features for Pixel-wise Video Correspondences
Rui Li, Shenglong Zhou, Dong Liu
Abstract
Video analysis tasks rely heavily on identifying the pixels from different frames that correspond to the same visual target. To tackle this problem, recent studies have advocated feature learning methods that aim to learn distinctive representations to match the pixels, especially in a self-supervised fashion. Unfortunately, these methods have difficulties for tiny or even single-pixel visual targets. Pixel-wise video correspondences were traditionally related to optical flows, which however lead to deterministic correspondences and lack robustness on real-world videos. We address the problem of learning features for establishing pixel-wise correspondences. Motivated by optical flows as well as the self-supervised feature learning, we propose to use not only labeled synthetic videos but also unlabeled real-world videos for learning fine-grained representations in a holistic framework. We adopt an adversarial learning scheme to enhance the generalization ability of the learned features. Moreover, we design a coarse-to-fine framework to pursue high computational efficiency. Our experimental results on a series of correspondence-based tasks demonstrate that the proposed method outperforms state-of-the-art rivals in both accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf00553b-44e4-428b-a090-974a70d67fdaCited by top-tier papers1
Ask how each one uses itBuilds on13
- DetCo: Unsupervised Contrastive Learning for Object DetectionEnze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan et al.ICCV 2021 · 364 citations
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 356 citations
- Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity PerspectiveJiarui Xu, Xiaolong WangICCV 2021 · 112 citations
- Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence LearningLiulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang et al.CVPR 2022 · 41 citations
- Dense Unsupervised Learning for Video SegmentationNikita Araslanov, Simone Schaub-Meyer, Stefan RothNeurIPS 2021 · 41 citations
Related papers
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 38 citations
- Self-supervised AutoFlowHsin-Ping Huang, Charles Herrmann, Junhwa Hur, Erika Lu et al.CVPR 2023
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
- Unified Mask Embedding and Correspondence Learning for Self-Supervised Video SegmentationLiulei Li, Wenguan Wang, Tianfei Zhou, Jianwu Li et al.CVPR 2023
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 23 citations
