Mitigating Semantic Collapse in Partially Relevant Video Retrieval
WonJun Moon, Minseok Jung, Gilhan Park, Tae-Young Kim, Cheol-Ho Cho, Woojin Jun, Jae-Pil Heo
Abstract
Partially Relevant Video Retrieval (PRVR) seeks videos where only part of the content matches a text query. Existing methods treat every annotated text-video pair as a positive and all others as negatives, ignoring the rich semantic variation both within a single video and across different videos. Consequently, embeddings of both queries and their corresponding video-clip segments for distinct events within the same video collapse together, while embeddings of semantically similar queries and segments from different videos are driven apart. This limits retrieval performance when videos contain multiple, diverse events. This paper addresses the aforementioned problems, termed as semantic collapse, in both the text and video embedding spaces. We first introduce Text Correlation Preservation Learning, which preserves the semantic relationships encoded by the foundation model across text queries. To address collapse in video embeddings, we propose Cross-Branch Video Alignment (CBVA), a contrastive alignment method that disentangles hierarchical video representations across temporal scales. Subsequently, we introduce order-preserving token merging and adaptive CBVA to enhance alignment by producing video segments that are internally coherent yet mutually distinctive. Extensive experiments on PRVR benchmarks demonstrate that our framework effectively prevents semantic collapse and substantially improves retrieval accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dcd311f3-5919-4b8b-be8a-7c85f6b01a12Cited by top-tier papers2
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
- Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video RetrievalJun Li, Peifeng Lai, Xuhang Lou, Jinpeng Wang et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video RetrievalCheol-Ho Cho, WonJun Moon, Woojin Jun, Minseok Jung et al.AAAI 2025 · 11 citations
- Action-and-object Aware Alignment for Partially Relevant Video RetrievalChuanshen Chen, Kai Zhou, Zhiquan Wen, Zeng You et al.AAAI 2026
- Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalWonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim et al.ICCV 2025 · 3 citations
- Bridging the Semantic Granularity Gap Between Text and Frame Representations for Partially Relevant Video RetrievalWoojin Jun, WonJun Moon, Cheol-Ho Cho, Minseok Jung et al.AAAI 2025 · 9 citations
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang et al.NeurIPS 2022 · 52 citations
