Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning
Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang, Jianwu Li, Yi Yang
摘要
Our target is to learn visual correspondence from unlabeled videos. We develop Liir, a locality-aware inter-and intra-video reconstruction method that fills in three missing pieces, i.e., instance discrimination, location awareness, and spatial compactness, of self-supervised correspondence learning puzzle. First, instead of most existing efforts focusing on intra-video self-supervision only, we exploit cross-video affinities as extra negative samples within a unified, inter-and intra-video reconstruction scheme. This enables instance discriminative representation learning by contrasting desired intra-video pixel association against negative inter-video correspondence. Second, we merge position information into correspondence matching, and design a position shifting strategy to remove the side-effect of position encoding during inter-video affinity computation, making our Liir location-sensitive. Third, to make full use of the spatial continuity nature of video data, we impose a compactness-based constraint on correspondence matching, yielding more sparse and reliable solutions. The learned representation surpasses self-supervised state-of-the-arts on label propagation tasks including objects, semantic parts, and keypoints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Exposing the Self-Supervised Space-Time Correspondence Learning via Graph KernelsZheyun Qin, Xiankai Lu, Xiushan Nie, Yilong Yin 等AAAI 2023 · 被引用 22 次
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 被引用 19 次
- VideoMAC: Video Masked Autoencoders Meet ConvNetsGensheng Pei, Tao Chen, Xiruo Jiang, Huafeng Liu 等CVPR 2024 · 被引用 14 次
- Learning Fine-Grained Features for Pixel-wise Video CorrespondencesRui Li, Shenglong Zhou, Dong LiuICCV 2023 · 被引用 7 次
- From ViT Features to Training-free Video Object Segmentation via Streaming-data Mixture ModelsRoy Uziel, Or Dinari, Oren FreifeldNeurIPS 2023 · 被引用 6 次
它引用的顶会 Paper32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- SOLOv2: Dynamic and Fast Instance SegmentationXinlong Wang, Rufeng Zhang, Tao Kong, Lei Li 等NeurIPS 2020 · 被引用 1,193 次
相关 Paper
- Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningZixu Zhao, Yueming Jin, Pheng-Ann HengICCV 2021 · 被引用 23 次
- Contrastive Transformation for Self-supervised Correspondence LearningNing Wang, Wengang Zhou, Houqiang LiAAAI 2021 · 被引用 38 次
- Space-Time Correspondence as a Contrastive Random WalkAllan Jabri, Andrew Owens, Alexei A. EfrosNeurIPS 2020 · 被引用 356 次
- Spatial-then-Temporal Self-Supervised Learning for Video CorrespondenceRui Li, Dong LiuCVPR 2023
- Contrastive Learning for Space-time Correspondence via Self-cycle ConsistencyJeany SonCVPR 2022 · 被引用 13 次
