Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
Yang Liu, Qianqian Xu, Peisong Wen, Siran Dai, Qingming Huang
Abstract
The rapid growth of online video resources has significantly promoted the development of video retrieval methods. As a standard evaluation metric for video retrieval, Average Precision (AP) assesses the overall rankings of relevant videos at the top list, making the predicted scores a reliable reference for the users. However, recent video retrieval methods utilize pair-wise losses that treat all sample pairs equally, leading to an evident gap between the training objective and evaluation metric. To effectively bridge this gap, in this work, we aim to address two primary challenges: a) The current similarity measure and AP-based loss are suboptimal for video retrieval; b) The noticeable noise from frame-to-frame matching introduces ambiguity in estimating the AP loss. In response to these challenges, we propose the Hierarchical learning framework for Average-Precision-oriented Video Retrieval (HAP-VR). For the former challenge, we develop the TopK-Chamfer Similarity and QuadLinear-AP loss to measure and optimize video-level similarities in terms of AP. For the latter challenge, we suggest constraining the frame-level similarities to achieve an accurate AP loss estimation. Experimental results present that HAP-VR outperforms existing methods on several benchmark datasets, providing a feasible solution for video retrieval tasks and thus offering potential benefits for the multi-media application.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55ced9b2-91e0-426c-a9ba-bd16d82fe277Cited by top-tier papers7
- LightFair: Towards an Efficient Alternative for Fair T2I Diffusion via Debiasing Pre-trained Text EncodersBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.NeurIPS 2025 · 17 citations
- Exploring Structural Degradation in Dense Representations for Self-supervised LearningSiran Dai, Qianqian Xu, Peisong Wen, Yang Liu et al.NeurIPS 2025 · 5 citations
- An Evaluation of N-Gram Selection Strategies for Regular Expression Indexing in Contemporary Text Analysis TasksLing Zhang, Shaleen Deep, Jignesh M. Patel, Karthikeyan SankaralingamVLDB 2025 · 2 citations
- Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual RepresentationBoyu Han, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.CVPR 2026 · 2 citations
- From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer LearningYang Liu, Qianqian Xu, Peisong Wen, Siran Dai et al.CVPR 2026 · 2 citations
Builds on19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- VmAP: A Fair Metric for Video Object DetectionAnupam Sobti, Vaibhav Mavi, M. Balakrishnan, Chetan AroraACM MM 2021 · 5 citations
- ViSiL: Fine-Grained Spatio-Temporal Video Similarity LearningGiorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Yiannis KompatsiarisICCV 2019 · 91 citations
- Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video RetrievalCheol-Ho Cho, WonJun Moon, Woojin Jun, Minseok Jung et al.AAAI 2025 · 11 citations
- PR-Net: Preference Reasoning for Personalized Video Highlight DetectionRunnan Chen, Penghao Zhou, Wenzhe Wang, Nenglun Chen et al.ICCV 2021 · 14 citations
- Learning With Average Precision: Training Image Retrieval With a Listwise LossJérôme Revaud, Jon Almazán, Rafael S. Rezende, César Roberto de SouzaICCV 2019 · 424 citations
