HANet: Hierarchical Alignment Networks for Video-Text Retrieval
Peng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv, Jing Liu
摘要
Video-text retrieval is an important yet challenging task in vision-language understanding, which aims to learn a joint embedding space where related video and text instances are close to each other. Most current works simply measure the video-text similarity based on video-level and text-level embeddings. However, the neglect of more fine-grained or local information causes the problem of insufficient representation. Some works exploit the local details by disentangling sentences, but overlook the corresponding videos, causing the asymmetry of video-text representation. To address the above limitations, we propose a Hierarchical Alignment Network (HANet) to align different level representations for video-text matching. Specifically, we first decompose video and text into three semantic levels, namely event (video and text), action (motion and verb), and entity (appearance and noun). Based on these, we naturally construct hierarchical representations in the individual-local-global manner, where the individual level focuses on the alignment between frame and word, local level focuses on the alignment between video clip and textual context, and global level focuses on the alignment between the whole video and text. Different level alignments capture fine-to-coarse correlations between video and text, as well as take the advantage of the complementary information among three semantic levels. Besides, our HANet is also richly interpretable by explicitly learning key semantic concepts. Extensive experiments on two public datasets, namely MSR-VTT and VATEX, show the proposed HANet outperforms other state-of-the-art methods, which demonstrates the effectiveness of hierarchical representation and alignment. Our code is publicly available at https://github.com/Roc-Ng/HANet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen 等ICCV 2023 · 被引用 35 次
- Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long VideosYulin Pan, Xiangteng He, Biao Gong, Yiliang Lv 等ICCV 2023 · 被引用 29 次
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang 等ACM MM 2022 · 被引用 26 次
它引用的顶会 Paper14
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze 等ICLR 2021 · 被引用 269 次
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 被引用 186 次
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 被引用 185 次
相关 Paper
- Fine-Grained Video-Text Retrieval With Hierarchical Graph ReasoningShizhe Chen, Yida Zhao, Qin Jin, Qi WuCVPR 2020
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng 等CVPR 2023
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 被引用 7 次
- Boosting Video-Text Retrieval with Explicit High-Level SemanticsHaoran Wang, Di Xu, Dongliang He, Fu Li 等ACM MM 2022 · 被引用 12 次
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang 等ACM MM 2021 · 被引用 47 次
