CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition Transformer
Ran Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan, Jie Zou, Yang Yang, Heng Tao Shen
Abstract
Video Moment Retrieval (VMR) involves locating specific moments within a video based on natural language queries. However, existing VMR methods that employ various strategies for cross-modal alignment still face challenges such as limited understanding of fine-grained semantics, semantic overlap, and sparse constraints. To address these limitations, we propose a novel Concept Decomposition Transformer (CDTR) model for VMR. CDTR introduces a semantic concept decomposition module that disentangles video moments and sentence queries into concept representations, reflecting the relevance between various concepts and capturing fine-grained semantics which is crucial for cross-modal matching. These decomposed concept representations are then used as pseudo-labels, determined as positive or negative samples by adaptive concept-specific thresholds. Subsequently, fine-grained concept alignment is performed in video intra-modal and textual-visual cross-modal, aligning different conceptual components within features, enhancing the model's ability to distinguish fine-grained semantics, and alleviating issues related to semantic overlap and sparse constraints. Comprehensive experiments demonstrate the effectiveness of the CDTR, outperforming state-of-the-art methods on three widely used datasets: QVHighlights, Charades-STA, and TACoS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46623bbf-53d0-4ba0-84e7-53599c20466bCited by top-tier papers2
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma et al.ICCV 2025 · 4 citations
- CVA: Context-aware Video-text Alignment for Video Temporal GroundingSungho Moon, Seunghun Lee, Jiwan Seo, Sunghoon ImCVPR 2026 · 4 citations
Related papers
- Maskable Retentive Network for Video Moment RetrievalJingjing Hu, Dan Guo, Kun Li, Zhan Si et al.ACM MM 2024 · 7 citations
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment RetrievalMingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong KimAAAI 2026 · 1 citation
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim et al.NeurIPS 2025 · 4 citations
