DPDV: Dual-Pathway and Dual-View Representation Learning for Bridging Information Asymmetry in Text-Video Retrieval
Zequn Xie, Xin Liu, Fangming Feng, Boyun Zhang, Tao Jin
Abstract
In recent years, CLIP-based text-video retrieval methods have developed rapidly, with research focusing on constructing diverse features and achieving effective interactions. However, the asymmetry of cross-modal information poses a challenge to accurately establishing retrieval relationships. To overcome this challenge, we propose a novel video retrieval framework, termed the Dual-Pathway and Dual-View model (DPDV), which consists of the Dual-Pathway Partitioning Module (DPPM) for constructing features at an appropriate granularity and the Dual-View Interaction Module (DVIM) for performing effective feature interactions. For DPPM, we simulate a human macro-level cognitive perspective by partitioning visual features into two categories based on their relevance to the text query and supplementing less relevant features with additional textual information. For DVIM, we simulate a human alignment strategy from macro to micro levels, focusing on local visual features while comprehensively modeling fine-grained interactions. We evaluate DPDV on five benchmark datasets, including MSRVTT, and achieve stateof-the-art performance on video retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
- Support-set bottlenecks for video-text representation learningMandela Patrick, Po-Yao Huang, Yuki Markus Asano, Florian Metze et al.ICLR 2021 · 269 citations
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan et al.CVPR 2022 · 190 citations
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian et al.CVPR 2022 · 132 citations
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou et al.ICCV 2023 · 98 citations
Related papers
- Relation Triplet Construction for Cross-modal Text-to-Video RetrievalXue Song, Jingjing Chen, Yu-Gang JiangACM MM 2023 · 8 citations
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsChan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang et al.CVPR 2025
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 53 citations
- Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningKaibin Tian, Yanhua Cheng, Yi Liu, Xinglin Hou et al.AAAI 2024 · 19 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
