Joint Searching and Grounding: Multi-Granularity Video Content Retrieval
Zhiguo Chen, Xun Jiang, Xing Xu, Zuo Cao, Yijun Mo, Heng Tao Shen
Abstract
Text-based video retrieval is a well-studied task aimed at retrieving relevant videos from a large collection in response to a given text query. Most existing TVR works assume that videos are already trimmed and fully relevant to the query thus ignoring that most videos in real-world scenarios are untrimmed and contain massive irrelevant video content. Moreover, as users' queries are only relevant to video events rather than complete videos, it is also more practical to provide specific video events rather than an untrimmed video list. In this paper, we introduce a challenging but more realistic task called Multi-Granularity Video Content Retrieval (MGVCR), which involves retrieving both video files and specific video content with their temporal locations. This task presents significant challenges since it requires identifying and ranking the partial relevance between long videos and text queries under the lack of temporal alignment supervision between the query and relevant moments. To this end, we propose a novel unified framework, termed, Joint Searching and Grounding (JSG). It consists of two branches: (1) a glance branch that coarsely aligns the query and moment proposals using inter-video contrastive learning, and (2) a gaze branch that finely aligns two modalities using both inter- and intra-video contrastive learning. Based on the glance-to-gaze design, our JSG method learns two separate joint embedding spaces for moments and text queries using a hybrid synergistic contrastive learning strategy. Extensive experiments on three public benchmarks, i.e., Charades-STA, DiDeMo, and ActivityNet-Captions demonstrate the superior performance of our JSG method on both video-level retrieval and event-level retrieval subtasks. Our open-source implementation code is available at https://github.com/CFM-MSG/Code_JSG.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 154e0026-4532-4e39-aff4-d55ce667bb68Cited by top-tier papers8
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and GranularitiesWoongyeong Yeo, Kangsan Kim, Soyeong Jeong, Jinheon Baek et al.ACL 2026 · 14 citations
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian et al.ICCV 2025 · 5 citations
- Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text RetrievalYang Du, Yuqi Liu, Qin JinACM MM 2024 · 4 citations
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
Related papers
- Video Corpus Moment Retrieval with Contrastive LearningHao Zhang, Aixin Sun, Wei Jing, Guoshun Nan et al.SIGIR 2021 · 88 citations
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 21 citations
- MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic LearningHongxu Ma, Guanshuo Wang, Fufu Yu, Qiong Jia et al.ACM MM 2025 · 9 citations
- GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment RetrievalMingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong KimAAAI 2026 · 1 citation
- Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight DetectionYicheng Xiao, Zhuoyan Luo, Yong Liu, Yue Ma et al.CVPR 2024 · 43 citations
