Learn from Unlabeled Videos for Near-duplicate Video Retrieval
Xiangteng He, Yulin Pan, Mingqian Tang, Yiliang Lv, Yuxin Peng
Abstract
Near-duplicate video retrieval (NDVR) aims to find the copies or transformations of the query video from a massive video database. It plays an important role in many video related applications, including copyright protection, tracing, filtering and etc. Video representation and similarity search are crucial to any video retrieval system. To derive effective video representation, most video retrieval systems require a large amount of manually annotated data for training, making it costly inefficient. In addition, most retrieval systems are based on frame-level features for video similarity searching, making it expensive both storage wise and search wise. To address the above issues, we propose a video representation learning (VRL) approach to effectively address the above shortcomings. It first effectively learns video representation from unlabeled videos via contrastive learning to avoid the expensive cost of manual annotation. Then, it exploits transformer structure to aggregate frame-level features into clip-level to reduce both storage space and search complexity. It can learn the complementary and discriminative information from the interactions among clip frames, as well as acquire the frame permutation and missing invariant ability to support more flexible retrieval manners. Comprehensive experiments on two challenging near-duplicate video retrieval datasets, namely FIVR-200K and SVD, verify the effectiveness of our proposed VRL approach, which achieves the best performance of video retrieval on accuracy and efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9e0c01f1-dad2-4575-9366-184a416da535Cited by top-tier papers3
- Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual RetrievalZhe Ma, Jianfeng Dong, Shouling Ji, Zhenguang Liu et al.AAAI 2024 · 14 citations
- Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video RetrievalYang Liu, Qianqian Xu, Peisong Wen, Siran Dai et al.ACM MM 2024 · 9 citations
- Tracing Copied Pixels and Regularizing Patch Affinity in Copy DetectionYichen Lu, Siwei Nie, Minlong Lu, Xudong Yang et al.ICCV 2025
Related papers
- SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video RetrievalQing-Yuan Jiang, Yi He, Gen Li, Jian Lin et al.ICCV 2019 · 52 citations
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga et al.NeurIPS 2020 · 59 citations
- Spatiotemporal Contrastive Video Representation LearningRui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang et al.CVPR 2021
- VVS: Video-to-Video Retrieval with Irrelevant Frame SuppressionWon Jo, Geuntaek Lim, Gwangjin Lee, Hyunwoo Kim et al.AAAI 2024 · 10 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
