Bridging the Semantic Granularity Gap Between Text and Frame Representations for Partially Relevant Video Retrieval
Woojin Jun, WonJun Moon, Cheol-Ho Cho, Minseok Jung, Jae-Pil Heo
摘要
Partially Relevant Video Retrieval (PRVR) addresses the challenges of text-to-video retrieval in real-world scenarios where untrimmed videos are prevalent. Traditional PRVR methods encode videos at two feature scales: (1) frame-level to capture fine details, and (2) clip-level to recognize broader content. However, these approaches align both scales with a single sentence representation, leading to suboptimal performance. In particular, we point out the level mismatch in aligning frame-level video features with a sentence representation, as the entire meaning of a sentence contains broader and more diverse content than what frame-level features can encode. This misalignment causes frame-level features to capture broader contexts and overlook local fine details. To tackle this issue, we propose a framework that represents a sentence as a set of multiple components, where each component aligns with frame-level semantics. Specifically, we introduce Semantic-Decomposed Matching (SDM) to adjust the granularity of the text description to match them with frame-level video features. In addition to the matching process, we develop the Adaptive Local Aggregator (ALA) to enhance video encoding in capturing finer local details, ensuring precise text-video alignment at the frame level. ALA adaptively integrates multi-scale local details within short temporal spans obtained by enforcing a strict temporal aggregation range. Finally, we reinforce detailed encoding at the frame level with newly designed objectives for both modalities. Extensive experiments integrating our framework with existing clip branches demonstrate its effectiveness and applicability, highlighting significant improvements in PRVR performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mitigating Semantic Collapse in Partially Relevant Video RetrievalWonJun Moon, Minseok Jung, Gilhan Park, Tae-Young Kim 等NeurIPS 2025 · 被引用 7 次
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang 等CVPR 2026 · 被引用 3 次
- Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video RetrievalJun Li, Peifeng Lai, Xuhang Lou, Jinpeng Wang 等ICML 2026
- Action-and-object Aware Alignment for Partially Relevant Video RetrievalChuanshen Chen, Kai Zhou, Zhiquan Wen, Zeng You 等AAAI 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ICCV 2023 · 被引用 98 次
- DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelPeng Jin, Hao Li, Zesen Cheng, Kehan Li 等ICCV 2023 · 被引用 95 次
相关 Paper
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 等ACM MM 2022 · 被引用 65 次
- Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video RetrievalCheol-Ho Cho, WonJun Moon, Woojin Jun, Minseok Jung 等AAAI 2025 · 被引用 11 次
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv 等ACM MM 2021 · 被引用 62 次
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen 等ICCV 2023 · 被引用 35 次
- Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalWonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim 等ICCV 2025 · 被引用 3 次
