Rethinking Noisy Video-Text Retrieval via Relation-aware Alignment
Huakai Lai, Guoxin Xiong, Huayu Mai, Xiang Liu, Tianzhu Zhang
Abstract
Video-Text Retrieval (VTR) is a core task in multi-modal understanding, drawing growing attention from both academia and industry in recent years. While numerous VTR methods have achieved success, most of them assume accurate visual-text correspondences during training, which is difficult to ensure in practice due to ubiquitous noise, known as noisy correspondences (NC). In this paper, we rethink how to mitigate the NC from the perspective of representative reference features (termed agents), and propose a novel relation-aware purified consistency (RPC) network to amend direct pairwise correlation, including representative agents construction and relation-aware ranking distribution alignment. The proposed RPC enjoys several merits. First, to learn the agents well without any correspondence supervision, we customize the agents construction according to the three characteristics of reliability, representativeness, and resilience. Second, the ranking distribution-based alignment process leverages the structural information inherent in inter-pair relationships, making it more robust compared to individual comparisons. Extensive experiments on five datasets under different settings demonstrate the efficacy and robustness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Two Losses, One Goal: Balancing Conflict Gradients for Semi-Supervised Semantic SegmentationRui Sun, Huayu Mai, Wangkai Li, Yujia Chen et al.ICCV 2025 · 1 citation
- Zero-Shot Multimodal Retrieval with Multi-Scale Contextual RepresentationsSourajit Saha, Tejas GokhaleACL 2026
- Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query ShiftsBingqing Zhang, Zhuo Cao, Heming Du, Yang Li et al.ICLR 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding et al.NeurIPS 2021 · 215 citations
Related papers
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence LearningZhuoyao Liu, Yang Liu, Wentao Feng, Shudong HuangAAAI 2026
- Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person RetrievalXinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang et al.AAAI 2026
- TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text RetrievalShuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu et al.AAAI 2025 · 2 citations
- Text Proxy: Decomposing Retrieval from a 1-to-N Relationship into N 1-to-1 Relationships for Text-Video RetrievalJian Xiao, Zhenzhen Hu, Jia Li, Richang HongAAAI 2025 · 7 citations
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang et al.NeurIPS 2022 · 52 citations
