CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition Transformer
Ran Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan, Jie Zou, Yang Yang, Heng Tao Shen
摘要
Video Moment Retrieval (VMR) involves locating specific moments within a video based on natural language queries. However, existing VMR methods that employ various strategies for cross-modal alignment still face challenges such as limited understanding of fine-grained semantics, semantic overlap, and sparse constraints. To address these limitations, we propose a novel Concept Decomposition Transformer (CDTR) model for VMR. CDTR introduces a semantic concept decomposition module that disentangles video moments and sentence queries into concept representations, reflecting the relevance between various concepts and capturing fine-grained semantics which is crucial for cross-modal matching. These decomposed concept representations are then used as pseudo-labels, determined as positive or negative samples by adaptive concept-specific thresholds. Subsequently, fine-grained concept alignment is performed in video intra-modal and textual-visual cross-modal, aligning different conceptual components within features, enhancing the model's ability to distinguish fine-grained semantics, and alleviating issues related to semantic overlap and sparse constraints. Comprehensive experiments demonstrate the effectiveness of the CDTR, outperforming state-of-the-art methods on three widely used datasets: QVHighlights, Charades-STA, and TACoS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma 等ICCV 2025 · 被引用 4 次
- CVA: Context-aware Video-text Alignment for Video Temporal GroundingSungho Moon, Seunghun Lee, Jiwan Seo, Sunghoon ImCVPR 2026 · 被引用 4 次
相关 Paper
- Maskable Retentive Network for Video Moment RetrievalJingjing Hu, Dan Guo, Kun Li, Zhan Si 等ACM MM 2024 · 被引用 7 次
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen 等CVPR 2022 · 被引用 150 次
- TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight DetectionHao Sun, Mingyao Zhou, Wenjing Chen, Wei XieAAAI 2024
- GranAlign: Granularity-Aware Alignment Framework for Zero-shot Video Moment RetrievalMingyu Jeon, Sunjae Yoon, Jonghee Kim, Junyeong KimAAAI 2026 · 被引用 1 次
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim 等NeurIPS 2025 · 被引用 4 次
