Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
Chen Jiang, Kaiming Huang, Sifeng He, Xudong Yang, Wei Zhang, Xiaobo Zhang, Yuan Cheng, Lei Yang, Qing Wang, Furong Xu, Tan Pan, Wei Chu
摘要
With the explosive growth of web videos in recent years, large-scale Content-Based Video Retrieval (CBVR) becomes increasingly essential in video filtering, recommendation, and copyright protection. Segment-level CBVR (S-CBVR) locates the start and end time of similar segments in finer granularity, which is beneficial for user browsing efficiency and infringement detection especially in long video scenarios. The challenge of S-CBVR task is how to achieve high temporal alignment accuracy with efficient computation and low storage consumption. In this paper, we propose a Segment Similarity and Alignment Network (SSAN) in dealing with the challenge which is firstly trained end-to-end in S-CBVR. SSAN is based on two newly proposed modules in video retrieval: (1) An efficient Self-supervised Keyframe Extraction (SKE) module to reduce redundant frame features, (2) A robust Similarity Pattern Detection (SPD) module for temporal alignment. In comparison with uniform frame extraction, SKE not only saves feature storage and search time, but also introduces comparable accuracy and limited extra computation time. In terms of temporal alignment, SPD localizes similar segments with higher accuracy and efficiency than existing deep learning methods. Furthermore, we jointly train SSAN with SKE and SPD and achieve an end-to-end improvement. Meanwhile, the two key modules SKE and SPD can also be effectively inserted into other video retrieval pipelines and gain considerable performance improvements. Experimental results on public datasets show that SSAN can obtain higher alignment accuracy while saving storage and online query computational cost compared to existing methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang 等AAAI 2023 · 被引用 26 次
- A Large-scale Comprehensive Dataset and Copy-overlap Aware Evaluation Protocol for Segment-level Video Copy DetectionSifeng He, Xudong Yang, Chen Jiang, Gang Liang 等CVPR 2022 · 被引用 18 次
- Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video CaptioningZeyu Xi, Haoying Sun, Yaofei Wu, Junchi Yan 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper4
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao 等AAAI 2020 · 被引用 182 次
- ViSiL: Fine-Grained Spatio-Temporal Video Similarity LearningGiorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Yiannis KompatsiarisICCV 2019 · 被引用 91 次
- SVD: A Large-Scale Short Video Dataset for Near-Duplicate Video RetrievalQing-Yuan Jiang, Yi He, Gen Li, Jian Lin 等ICCV 2019 · 被引用 52 次
- Metric Learning with Equidistant and Equidistributed Triplet-based Loss for Product Image SearchFurong Xu, Wei Zhang, Yuan Cheng, Wei ChuWWW 2020 · 被引用 13 次
相关 Paper
- VVS: Video-to-Video Retrieval with Irrelevant Frame SuppressionWon Jo, Geuntaek Lim, Gwangjin Lee, Hyunwoo Kim 等AAAI 2024 · 被引用 10 次
- Video Similarity and Alignment Learning on Partial Video Copy DetectionZhen Han, Xiangteng He, Mingqian Tang, Yiliang LvACM MM 2021 · 被引用 34 次
- Learn from Unlabeled Videos for Near-duplicate Video RetrievalXiangteng He, Yulin Pan, Mingqian Tang, Yiliang Lv 等SIGIR 2022 · 被引用 21 次
- Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity PerspectiveJiarui Xu, Xiaolong WangICCV 2021 · 被引用 112 次
- Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosYubin Hu, Yuze He, Yanghao Li, Jisheng Li 等CVPR 2023
