Relation Triplet Construction for Cross-modal Text-to-Video Retrieval
Xue Song, Jingjing Chen, Yu-Gang Jiang
Abstract
Cross-modal text-to-video retrieval aims to find semantically related videos for a text query. Since video and text are distinct modalities, the major challenge comes from building the correspondence between two modalities, thus relevant samples could be matched. Inherently, the text contains multiple relatively complete semantic units and each one is composed of three primary components, i.e., subject, predicate and object (SVO triplet). Therefore, it requires similar modeling of video content -- objects and their relations, to correctly retrieve videos for texts. To model fine-grained visual relations, this paper proposes a Multi-Granularity Matching (MGM) framework that considers both fine-grained relation triplet matching and coarse-grained global semantic matching for text-to-video retrieval. Specifically, in the proposed framework, we represent videos as SVO triplet tracklets by extracting frame-level relation triplets followed by temporal relation association across frames. Moreover, we design a transformer-based Bi-directional Fusion Block (BFB) to express each SVO triplet with a highly unified representation. The constructed SVO triplet tracklets provide a reasonable way to model fine-grained video contents, fulfilling a better alignment between videos and texts. Extensive experiments conducted on three benchmark datasets, i.e., MSR-VTT, LSMDC and MSVD, demonstrate the effectiveness of our proposed method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 09bdaa2c-b879-4900-ba62-6c572302f5f9Cited by top-tier papers6
- Exposing the Deception: Uncovering More Forgery Clues for Deepfake DetectionZhongjie Ba, Qingyu Liu, Zhenguang Liu, Shuang Wu et al.AAAI 2024 · 101 citations
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 4 citations
- Asymmetric Cross-Modal Hashing Based on Formal Concept AnalysisYinan Li, Jun Long, Zhan YangAAAI 2025 · 4 citations
- Open-Vocabulary Video Relation ExtractionWentao Tian, Zheng Wang, Yuqian Fu, Jingjing Chen et al.AAAI 2024 · 2 citations
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu et al.SIGIR 2026
Related papers
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Unified Coarse-to-Fine Alignment for Video-Text RetrievalZiyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius et al.ICCV 2023 · 90 citations
- Progressive Semantic Matching for Video-Text RetrievalHongying Liu, Ruyi Luo, Fanhua Shang, Mantang Niu et al.ACM MM 2021 · 20 citations
- MPT: Multi-grained Prompt Tuning for Text-Video RetrievalHaonan Zhang, Pengpeng Zeng, Lianli Gao, Jingkuan Song et al.ACM MM 2024 · 16 citations
- DPDV: Dual-Pathway and Dual-View Representation Learning for Bridging Information Asymmetry in Text-Video RetrievalZequn Xie, Xin Liu, Fangming Feng, Boyun Zhang et al.ACL 2026
