Compositional Transformation Reasoning for Composed Video Retrieval
Sihong Huang, Jiaxin Wu, Dongmei Jiang, Yi Cai, Yaowei Wang, Xiaoyong Wei
摘要
Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities in response to fine-grained textual edits. Existing methods address this issue by training on large-scale video-text-video triplets or by generating dense textual descriptions to capture subtle visual differences. However, these supervised approaches often rely on noisy web-scale data and dataset-specific correspondences, leading to overfitting and limited generalization in diverse or fine-grained scenarios, while also failing to effectively model compositional and temporal transformations. We propose Multiobjective Reasoning (MoRe), a zero-shot framework based on MLLMs for multi-objective candidate selection and finegrained transformation reasoning. Our method decomposes the compositional transformation into three complementary reasoning dimensions, i.e., entity, action, and scene, and performs pairwise candidate reasoning to explicitly capture semantic evolution over time. Furthermore, we introduce a recall-oriented multi-objective candidate selection module that identifies high-quality retrieval targets by jointly balancing visual, textual, and multimodal similarities before transformation reasoning. Experiments on EgoCVR and WebVid-CoVR demonstrate the effectiveness of our method over state-of-the-art approaches under the zero-shot setting, with R@1 improvements of +5.8 and +10.8, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
相关 Paper
- Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsOmkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer, Salman H. Khan 等CVPR 2024 · 被引用 10 次
- Beyond Simple Edits: Composed Video Retrieval with Dense ModificationsOmkar Thawakar, Dmitry Demidov, Ritesh Thawkar, Rao Muhammad Anwer 等ICCV 2025 · 被引用 2 次
- CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image RetrievalZelong Sun, Dong Jing, Zhiwu LuICCV 2025 · 被引用 5 次
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang 等AAAI 2026 · 被引用 24 次
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang 等CVPR 2026 · 被引用 3 次
