SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video Retrieval
Qing-Yuan Jiang, Yi He, Gen Li, Jian Lin, Lei Li, Wu-Jun Li
Abstract
With the explosive growth of video data in real applications, near-duplicate video retrieval (NDVR) has become indispensable and challenging, especially for short videos. However, all existing NDVR datasets are introduced for long videos. Furthermore, most of them are small-scale and lack of diversity due to the high cost of collecting and labeling near-duplicate videos. In this paper, we introduce a large-scale short video dataset, called SVD, for the ND-VR task. SVD contains over 500,000 short videos and over 30,000 labeled videos of near-duplicates. We use multiple video mining techniques to construct positive/negative pairs. Furthermore, we design temporal and spatial transformations to mimic user-attack behavior in real applications for constructing more difficult variants of SVD. Experiments show that existing state-of-the-art NDVR methods, including real-value based and hashing based methods, fail to achieve satisfactory performance on this challenging dataset. The release of SVD dataset will foster research and system engineering in the NDVR area. The SVD dataset is available at https://svdbase.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7960b693-6a87-4abd-b6e6-37b3f292e6ebCited by top-tier papers9
- Aligning Distillation For Cold-start Item RecommendationFeiran Huang, Zefan Wang, Xiao Huang, Yufeng Qian et al.SIGIR 2023 · 100 citations
- Learning Segment Similarity and Alignment in Large-Scale Content Based Video RetrievalChen Jiang, Kaiming Huang, Sifeng He, Xudong Yang et al.ACM MM 2021 · 35 citations
- Video Similarity and Alignment Learning on Partial Video Copy DetectionZhen Han, Xiangteng He, Mingqian Tang, Yiliang LvACM MM 2021 · 34 citations
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang et al.AAAI 2023 · 26 citations
- A Large-scale Comprehensive Dataset and Copy-overlap Aware Evaluation Protocol for Segment-level Video Copy DetectionSifeng He, Xudong Yang, Chen Jiang, Gang Liang et al.CVPR 2022 · 18 citations
Related papers
- Learn from Unlabeled Videos for Near-duplicate Video RetrievalXiangteng He, Yulin Pan, Mingqian Tang, Yiliang Lv et al.SIGIR 2022 · 21 citations
- Spatiotemporal Fine-grained Video Description for Short VideosTe Yang, Jian Jia, Bo Wang, Yanhua Cheng et al.ACM MM 2024 · 1 citation
- KoDF: A Large-scale Korean DeepFake Detection DatasetPatrick Kwon, Jaeseong You, Gyuhyeon Nam, Sungwoo Park et al.ICCV 2021 · 154 citations
- Your One-Stop Solution for AI-Generated Video DetectionLong Ma, Zihao Xue, Yan Wang, Zhiyuan Yan et al.CVPR 2026 · 13 citations
- Celeb-DF: A Large-Scale Challenging Dataset for DeepFake ForensicsYuezun Li, Xin Yang, Pu Sun, Honggang Qi et al.CVPR 2020
