Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
Jun Li, Jinpeng Wang, Chaolei Tan, Niu Lian, Long Chen, Yaowei Wang, Min Zhang, Shu-Tao Xia, Bin Chen
Abstract
Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos and overlooks certain hierarchical semantics, ultimately leading to suboptimal temporal modeling. To address this issue, we propose the first hyperbolic modeling framework for PRVR, namely HLFormer, which leverages hyperbolic space learning to compensate for the suboptimal hierarchical modeling capabilities of Euclidean space. Specifically, HLFormer integrates the Lorentz Attention Block and Euclidean Attention Block to encode video embeddings in hybrid spaces, using the Mean-Guided Adaptive Interaction Module to dynamically fuse features. Additionally, we introduce a Partial Order Preservation Loss to enforce "text ≺ video" hierarchy through Lorentzian cone constraints. This approach further enhances cross-modal matching by reinforcing partial relevance between video content and text queries. Extensive experiments show that HLFormer outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICCV25-HLFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b9c45b5-c195-4295-86c5-2591dae49475Cited by top-tier papers5
- Path-Decoupled Hyperbolic Flow Matching for Few-Shot AdaptationLin Li, Ziqi Jiang, Gefan Ye, Zhenqi He et al.ICML 2026 · 4 citations
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
- Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video RetrievalJun Li, Peifeng Lai, Xuhang Lou, Jinpeng Wang et al.ICML 2026
- Attend to Anything: Foundation Model for Unified Human Attention ModelingWenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun ZhaoICML 2026
- Action-and-object Aware Alignment for Partially Relevant Video RetrievalChuanshen Chen, Kai Zhou, Zhiquan Wen, Zeng You et al.AAAI 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 791 citations
- Hyperbolic Image-text RepresentationsKaran Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson et al.ICML 2023 · 137 citations
- Hyperbolic Vision Transformers: Combining Improvements in Metric LearningAleksandr Ermolov, Leyla Mirvakhabova, Valentin Khrulkov, Nicu Sebe et al.CVPR 2022 · 97 citations
Related papers
- GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video RetrievalYuting Wang, Jinpeng Wang, Bin Chen, Ziyun Zeng et al.AAAI 2024 · 32 citations
- Mitigating Semantic Collapse in Partially Relevant Video RetrievalWonJun Moon, Minseok Jung, Gilhan Park, Tae-Young Kim et al.NeurIPS 2025 · 7 citations
- Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D RetrievalWenrui Li, Yidan Lu, Yeyu Chai, Rui Zhao et al.AAAI 2026
- Hyperbolic Gramian Volumes for Multimodal AlignmentSaiyang Na, Feng Jiang, Qifeng Zhou, Wenliang Zhong et al.CVPR 2026
- Fully Hyperbolic Convolutional Neural Networks for Computer VisionAhmad Bdeir, Kristian Schwethelm, Niels LandwehrICLR 2024 · 45 citations
