HANet: Hierarchical Alignment Networks for Video-Text Retrieval
Peng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv, Jing Liu
Abstract
Video-text retrieval is an important yet challenging task in vision-language understanding, which aims to learn a joint embedding space where related video and text instances are close to each other. Most current works simply measure the video-text similarity based on video-level and text-level embeddings. However, the neglect of more fine-grained or local information causes the problem of insufficient representation. Some works exploit the local details by disentangling sentences, but overlook the corresponding videos, causing the asymmetry of video-text representation. To address the above limitations, we propose a Hierarchical Alignment Network (HANet) to align different level representations for video-text matching. Specifically, we first decompose video and text into three semantic levels, namely event (video and text), action (motion and verb), and entity (appearance and noun). Based on these, we naturally construct hierarchical representations in the individual-local-global manner, where the individual level focuses on the alignment between frame and word, local level focuses on the alignment between video clip and textual context, and global level focuses on the alignment between the whole video and text. Different level alignments capture fine-to-coarse correlations between video and text, as well as take the advantage of the complementary information among three semantic levels. Besides, our HANet is also richly interpretable by explicitly learning key semantic concepts. Extensive experiments on two public datasets, namely MSR-VTT and VATEX, show the proposed HANet outperforms other state-of-the-art methods, which demonstrates the effectiveness of hierarchical representation and alignment. Our code is publicly available at https://github.com/Roc-Ng/HANet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20f4a839-e12c-48a8-b127-772d4ef3a60cCited by top-tier papers17
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen et al.ICCV 2023 · 35 citations
- Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long VideosYulin Pan, Xiangteng He, Biao Gong, Yiliang Lv et al.ICCV 2023 · 29 citations
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
Builds on14
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze et al.ICLR 2021 · 269 citations
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 186 citations
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 185 citations
Related papers
- Fine-Grained Video-Text Retrieval With Hierarchical Graph ReasoningShizhe Chen, Yida Zhao, Qin Jin, Qi WuCVPR 2020
- Hierarchical Semantic Correspondence Networks for Video Paragraph GroundingChaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng et al.CVPR 2023
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 7 citations
- Boosting Video-Text Retrieval with Explicit High-Level SemanticsHaoran Wang, Di Xu, Dongliang He, Fu Li et al.ACM MM 2022 · 12 citations
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang et al.ACM MM 2021 · 47 citations
