Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval
Yue Wu, Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang, Shuhui Wang
Abstract
With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual information. However, the inherent heterogeneity between the modalities poses significant challenges. Textual data are highly abstract, while video content contains substantial redundancy. The modality gap in information representation makes existing methods struggle with the modality fusion and alignment required for fine-grained composed retrieval. To overcome these challenges, we first introduce FineCVR-1M, a finegrained composed video retrieval dataset containing 1,010,071 video-text triplets with detailed textual descriptions. This dataset is constructed through an automated process that identifies key concept changes between video pairs to generate textual descriptions for both static and action concepts. For fine-grained retrieval methods, the key challenge lies in understanding the detailed requirements. Text description serves as clear expressions of intent, but it requires models to distinguish subtle differences in the description of video semantics. Therefore, we propose a textual Feature Disentanglement and Cross-modal Alignment framework (FDCA) that disentangles features at both the sentence and token levels. At the sequence level, we separate text features into retained and injected features. At the token level, an Auxiliary Token Disentangling mechanism is proposed to disentangle texts into retained, injected, and excluded tokens. The disentanglement at both levels extracts fine-grained features, which are aligned and fused with the reference video to extract global representations for video retrieval. Experiments on FineCVR-1M dataset demonstrate the superior performance of FDCA. Our code and dataset are available at: https://may2333.github.io/FineCVR/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1ed5d878-3e62-4e10-b276-d65a98160f64Cited by top-tier papers6
- ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Qinlei Huang et al.AAAI 2026 · 24 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu et al.ACM MM 2025 · 5 citations
- Compositional Transformation Reasoning for Composed Video RetrievalSihong Huang, Jiaxin Wu, Dongmei Jiang, Yi Cai et al.CVPR 2026 · 3 citations
- OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and TextJunyang Ji, Shengjun Zhang, Da Li, Yuxiao Luo et al.ICLR 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Beyond Simple Edits: Composed Video Retrieval with Dense ModificationsOmkar Thawakar, Dmitry Demidov, Ritesh Thawkar, Rao Muhammad Anwer et al.ICCV 2025 · 2 citations
- Fine-Grained Video-Text Retrieval With Hierarchical Graph ReasoningShizhe Chen, Yida Zhao, Qin Jin, Qi WuCVPR 2020
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang et al.ACM MM 2021 · 47 citations
- Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence LearningJuncheng Li, Junlin Xie, Long Qian, Linchao Zhu et al.CVPR 2022 · 63 citations
- CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition TransformerRan Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan et al.AAAI 2025 · 6 citations
