VVS: Video-to-Video Retrieval with Irrelevant Frame Suppression
Won Jo, Geuntaek Lim, Gwangjin Lee, Hyunwoo Kim, Byungsoo Ko, Yukyung Choi
Abstract
In content-based video retrieval (CBVR), dealing with large-scale collections, efficiency is as important as accuracy; thus, several video-level feature-based studies have actively been conducted. Nevertheless, owing to the severe difficulty of embedding a lengthy and untrimmed video into a single feature, these studies have been insufficient for accurate retrieval compared to frame-level feature-based studies. In this paper, we show that appropriate suppression of irrelevant frames can provide insight into the current obstacles of the video-level approaches. Furthermore, we propose a Video-to-Video Suppression network (VVS) as a solution. VVS is an end-to-end framework that consists of an easy distractor elimination stage to identify which frames to remove and a suppression weight generation stage to determine the extent to suppress the remaining frames. This structure is intended to effectively describe an untrimmed video with varying content and meaningless information. Its efficacy is proved via extensive experiments, and we show that our approach is not only state-of-the-art in video-level approaches but also has a fast inference time despite possessing retrieval capabilities close to those of frame-level approaches. Code is available at https://github.com/sejong-rcv/VVS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23cd5d10-0833-4a8a-8a00-a2a3cec38a8eCited by top-tier papers2
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 11 citations
- Collaborative Weakly Supervised Video Correlation Learning for Procedure-Aware Instructional Video AnalysisTianyao He, Huabin Liu, Yuxi Li, Xiao Ma et al.AAAI 2024 · 8 citations
Builds on4
- ViSiL: Fine-Grained Spatio-Temporal Video Similarity LearningGiorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Yiannis KompatsiarisICCV 2019 · 91 citations
- Symmetrical Synthesis for Deep Metric LearningGeonmo Gu, ByungSoo KoAAAI 2020 · 26 citations
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 26 citations
- Embedding Expansion: Augmentation in Embedding Space for Deep Metric LearningByungSoo Ko, Geonmo GuCVPR 2020
Related papers
- Learning Segment Similarity and Alignment in Large-Scale Content Based Video RetrievalChen Jiang, Kaiming Huang, Sifeng He, Xudong Yang et al.ACM MM 2021 · 35 citations
- Learn from Unlabeled Videos for Near-duplicate Video RetrievalXiangteng He, Yulin Pan, Mingqian Tang, Yiliang Lv et al.SIGIR 2022 · 21 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
- Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy ReductionChaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat et al.ICCV 2023 · 4 citations
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen et al.ICCV 2023 · 35 citations
