Fine-grained Cross-modal Alignment Network for Text-Video Retrieval
Ning Han, Jingjing Chen, Guangyi Xiao, Hao Zhang, Yawen Zeng, Hao Chen
Abstract
Despite the recent progress of cross-modal text-to-video retrieval techniques, their performance is still unsatisfactory. Most existing works follow a trend of learning a joint embedding space to measure the distance between global-level or local-level textual and video representation. The fine-grained interactions between video segments and phrases are usually neglected in cross-modal learning, which results in suboptimal retrieval performances. To tackle the problem, we propose a novel Fine-grained Cross-modal Alignment Network (FCA-Net), which considers the interactions between visual semantic units (i.e., sub-actions/sub-events) in videos and phrases in sentences for cross-modal alignment. Specifically, the interactions between visual semantic units and phrases are formulated as a link prediction problem optimized by a graph auto-encoder to obtain the explicit relations between them and enhance the aligned feature representation for fine-grained cross-modal alignment. Experimental results on MSR-VTT, YouCook2, and VATEX datasets demonstrate the superiority of our model as compared to the state-of-the-art method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers18
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 48 citations
- Multi-Modal Knowledge Hypergraph for Diverse Image RetrievalYawen Zeng, Qin Jin, Tengfei Bao, Wenfeng LiAAAI 2023 · 40 citations
- Modality-Independent Graph Neural Networks with Global Transformers for Multimodal RecommendationJun Hu, Bryan Hooi, Bingsheng He, Yinwei WeiAAAI 2025 · 31 citations
- Attacking Video Recognition Models with Bullet-Screen CommentsKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu et al.AAAI 2022 · 27 citations
Related papers
- EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video RetrievalYuhan Chen, Pengwen Dai, Chuan Wang, Dayan Wu et al.CVPR 2026
- Progressive Semantic Matching for Video-Text RetrievalHongying Liu, Ruyi Luo, Fanhua Shang, Mantang Niu et al.ACM MM 2021 · 20 citations
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 7 citations
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
- Learning Semantic Alignment with Global Modality Reconstruction for Video-Language Pre-training towards RetrievalMingchao Li, Xiaoming Shi, Haitao Leng, Wei Zhou et al.AAAI 2023 · 4 citations
