Lune

CVPR2026顶会

Compositional Transformation Reasoning for Composed Video Retrieval

Sihong Huang, Jiaxin Wu, Dongmei Jiang, Yi Cai, Yaowei Wang, Xiaoyong Wei

出版方
2026年份
3被引次数

摘要

Composed Video Retrieval aims to retrieve a target video given a reference video and a textual modification describing the desired change. The core challenge lies in modeling compositional multimodal transformations, i.e., how entities, actions, and scenes evolve across video and language modalities in response to fine-grained textual edits. Existing methods address this issue by training on large-scale video-text-video triplets or by generating dense textual descriptions to capture subtle visual differences. However, these supervised approaches often rely on noisy web-scale data and dataset-specific correspondences, leading to overfitting and limited generalization in diverse or fine-grained scenarios, while also failing to effectively model compositional and temporal transformations. We propose Multiobjective Reasoning (MoRe), a zero-shot framework based on MLLMs for multi-objective candidate selection and finegrained transformation reasoning. Our method decomposes the compositional transformation into three complementary reasoning dimensions, i.e., entity, action, and scene, and performs pairwise candidate reasoning to explicitly capture semantic evolution over time. Furthermore, we introduce a recall-oriented multi-objective candidate selection module that identifies high-quality retrieval targets by jointly balancing visual, textual, and multimodal similarities before transformation reasoning. Experiments on EgoCVR and WebVid-CoVR demonstrate the effectiveness of our method over state-of-the-art approaches under the zero-shot setting, with R@1 improvements of +5.8 and +10.8, respectively.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖