ViSiL: Fine-Grained Spatio-Temporal Video Similarity Learning
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Yiannis Kompatsiaris
摘要
In this paper we introduce ViSiL, a Video Similarity Learning architecture that considers fine-grained Spatio-Temporal relations between pairs of videos -- such relations are typically lost in previous video retrieval approaches that embed the whole frame or even the whole video into a vector descriptor before the similarity estimation. By contrast, our Convolutional Neural Network (CNN)-based approach is trained to calculate video-to-video similarity from refined frame-to-frame similarity matrices, so as to consider both intra- and inter-frame relations. In the proposed method, pairwise frame similarity is estimated by applying Tensor Dot (TD) followed by Chamfer Similarity (CS) on regional CNN frame features - this avoids feature aggregation before the similarity calculation between frames. Subsequently, the similarity matrix between all video frames is fed to a four-layer CNN, and then summarized using Chamfer Similarity (CS) into a video-to-video similarity score -- this avoids feature aggregation before the similarity calculation between videos and captures the temporal similarity patterns between matching frame sequences. We train the proposed network using a triplet loss scheme and evaluate it on five public benchmark datasets on four different video retrieval problems where we demonstrate large improvements in comparison to the state of the art. The implementation of ViSiL is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Spatio-Temporal Trajectory Similarity Learning in Road NetworksZiquan Fang, Yuntao Du, Xinjun Zhu, Danlei Hu 等KDD 2022 · 被引用 68 次
- Learning Segment Similarity and Alignment in Large-Scale Content Based Video RetrievalChen Jiang, Kaiming Huang, Sifeng He, Xudong Yang 等ACM MM 2021 · 被引用 35 次
- Video Similarity and Alignment Learning on Partial Video Copy DetectionZhen Han, Xiangteng He, Mingqian Tang, Yiliang LvACM MM 2021 · 被引用 34 次
- TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible SupervisionSifeng He, Yue He, Minlong Lu, Chen Jiang 等AAAI 2023 · 被引用 26 次
- UBoCo: Unsupervised Boundary Contrastive Learning for Generic Event Boundary DetectionHyolim Kang, Jinwoo Kim, Taehyun Kim, Seon Joo KimCVPR 2022 · 被引用 26 次
相关 Paper
- Dense Interaction Learning for Video-based Person Re-identificationTianyu He, Xin Jin, Xu Shen, Jianqiang Huang 等ICCV 2021 · 被引用 73 次
- VVS: Video-to-Video Retrieval with Irrelevant Frame SuppressionWon Jo, Geuntaek Lim, Gwangjin Lee, Hyunwoo Kim 等AAAI 2024 · 被引用 10 次
- Self-supervised Video Representation Learning Using Inter-intra Contrastive FrameworkLi Tao, Xueting Wang, Toshihiko YamasakiACM MM 2020 · 被引用 110 次
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga 等NeurIPS 2020 · 被引用 59 次
- Rethinking Self-supervised Correspondence Learning: A Video Frame-level Similarity PerspectiveJiarui Xu, Xiaolong WangICCV 2021 · 被引用 112 次
