HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video Retrieval
Zhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu, Haokun Wen, Weili Guan
摘要
Composed Video Retrieval (CVR) is a challenging video retrieval task that utilizes multi-modal queries, consisting of a reference video and modification text, to retrieve the desired target video. The core of this task lies in understanding the multi-modal composed query and achieving accurate composed feature learning. Within multi-modal queries, the video modality typically carries richer semantic content compared to the textual modality. However, previous works have largely overlooked the disparity in information density between these two modalities. This limitation can lead to two critical issues: 1) modification subject referring ambiguity and 2) limited detailed semantic focus, both of which degrade the performance of CVR models. To address the aforementioned issues, we propose a novel CVR framework, namely the Hierarchical Uncertainty-aware Disambiguation network (HUD). HUD is the first framework that leverages the disparity in information density between video and text to enhance multi-modal query understanding. It comprises three key components: (a) Holistic Pronoun Disambiguation, (b) Atomistic Uncertainty Modeling, and (c) Holistic-to-Atomistic Alignment. By exploiting overlapping semantics through holistic cross-modal interaction and fine-grained semantic alignment via atomistic-level cross-modal interaction, HUD enables effective object disambiguation and enhances the focus on detailed semantics, thereby achieving precise composed feature learning. Moreover, our proposed HUD is also applicable to the Composed Image Retrieval (CIR) task and achieves state-of-the-art performance across three benchmark datasets for both CVR and CIR tasks. The codes are available on https://zivchen-ty.github.io/HUD.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SD-MVS: Segmentation-Driven Deformation Multi-View Stereo with Spherical Refinement and EM OptimizationZhenlong Yuan, Jiakai Cao, Zhaoxin Li, Hao Jiang 等AAAI 2024 · 被引用 38 次
- MSP-MVS: Multi-Granularity Segmentation Prior Guided Multi-View StereoZhenlong Yuan, Cong Liu, Fei Shen, Zhaoxin Li 等AAAI 2025 · 被引用 22 次
- DVP-MVS: Synergize Depth-Edge and Visibility Prior for Multi-View StereoZhenlong Yuan, Jinguo Luo, Fei Shen, Zhaoxin Li 等AAAI 2025 · 被引用 19 次
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang 等CVPR 2026 · 被引用 16 次
- Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image RetrievalZhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang 等CVPR 2026 · 被引用 16 次
它引用的顶会 Paper55
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 被引用 344 次
相关 Paper
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsOmkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer, Salman H. Khan 等CVPR 2024 · 被引用 10 次
- Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningHaomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng 等AAAI 2026 · 被引用 1 次
- Beyond Simple Edits: Composed Video Retrieval with Dense ModificationsOmkar Thawakar, Dmitry Demidov, Ritesh Thawkar, Rao Muhammad Anwer 等ICCV 2025 · 被引用 2 次
- Modeling Uncertainty in Composed Image Retrieval via Probabilistic EmbeddingsHaomiao Tang, Jinpeng Wang, Yuang Peng, Guanghao Meng 等ACL 2025 · 被引用 8 次
