Gravitation-Driven Semantic Alignment for Text Video Retrieval
Yi Yang, Zheng Wang, Xing Xu, Jingkuan Song, Heng Tao Shen
摘要
The inherent semantic ambiguity of "many-to-many", where one video matches multiple texts and vice versa, aggravates the difficulty in text-video retrieval. The dominant deterministic embeddings only struggle to capture the mean semantics, while existing probabilistic methods fail to distinguish hard negatives for their imposing rigid uncertainty priors or ignoring the interaction between similarity and uncertainty. To this end, we propose a novel physics-inspired framework (GraviAlign) that decomposes the alignment of cross-modal semantic distributions into two orthogonal factors inspired by the Gravitational Force:
(1) Semantic Attraction measuring gravitational alignment between distribution centers via uncertainty-derived "semantic mass" and "semantic distance"; (2) Geometric Overlap quantifying distribution intersection. Each factor has independent veto power to reject those matches with misalignment or poor overlap. Additionally, GraviAlign offers an efficient (O(D)), theoretically grounded alternative to intractable joint integrals. Extensive experiments on DiDeMo, MSR-VTT, and ActivityNet demonstrate our effectiveness and superiority, and solid ablation studies confirm the indispensability of two novel components.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan 等ACM MM 2022 · 被引用 314 次
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan 等CVPR 2022 · 被引用 190 次
- HiT: Hierarchical Transformer with Momentum Contrast for Video-Text RetrievalSong Liu, Haoqi Fan, Shengsheng Qian, Yiru Chen 等ICCV 2021 · 被引用 172 次
- CenterCLIP: Token Clustering for Efficient Text-Video RetrievalShuai Zhao, Linchao Zhu, Xiaohan Wang, Yi YangSIGIR 2022 · 被引用 150 次
相关 Paper
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang 等NeurIPS 2022 · 被引用 52 次
- T2VLAD: Global-Local Sequence Alignment for Text-Video RetrievalXiaohan Wang, Linchao Zhu, Yi YangCVPR 2021
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 被引用 26 次
- Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalXiaoshuai Hao, Wanqian Zhang, Dayan Wu, Fei Zhu 等CVPR 2023
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ICCV 2023 · 被引用 98 次
