Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval
Zhihang Liu, Jun Li, Hongtao Xie, Pandeng Li, Jiannan Ge, Sun'ao Liu, Guoqing Jin
摘要
Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic richness inherent in videos far exceeds that of a given limited-length sentence. Therefore, in pursuit of better alignment, a natural idea is enhancing the video modality to filter out query-irrelevant semantics, and enhancing the text modality to capture more segment-relevant knowledge. In this paper, we introduce Modal-Enhanced Semantic Modeling (MESM), a novel framework for more balanced alignment through enhancing features at two levels. First, we enhance the video modality at the frame-word level through word reconstruction. This strategy emphasizes the portions associated with query words in frame-level features while suppressing irrelevant parts. Therefore, the enhanced video contains less redundant semantics and is more balanced with the textual modality. Second, we enhance the textual modality at the segment-sentence level by learning complementary knowledge from context sentences and ground-truth segments. With the knowledge added to the query, the textual modality thus maintains more meaningful semantics and is more balanced with the video modality. By implementing two levels of MESM, the semantic information from both modalities is more balanced to align, thereby bridging the modality gap. Experiments on three widely used benchmarks, including the out-of-distribution settings, show that the proposed framework achieves a new start-of-the-art performance with notable generalization ability (e.g., 4.42% and 7.69% average gains of R1@0.7 on Charades-STA and Charades-CG). The code will be available at https://github.com/lntzm/MESM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingBoqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin WangCVPR 2024 · 被引用 9 次
- Multi-Scale Contrastive Learning for Video Temporal GroundingThong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu 等AAAI 2025 · 被引用 7 次
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong 等ICCV 2025 · 被引用 5 次
- Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalDezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin 等AAAI 2025 · 被引用 4 次
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu 等NeurIPS 2022 · 被引用 790 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
相关 Paper
- Learning Semantic Alignment with Global Modality Reconstruction for Video-Language Pre-training towards RetrievalMingchao Li, Xiaoming Shi, Haitao Leng, Wei Zhou 等AAAI 2023 · 被引用 4 次
- Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalDaizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan 等ACM MM 2023 · 被引用 10 次
- CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition TransformerRan Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan 等AAAI 2025 · 被引用 6 次
- Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-TrainingDezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin 等CVPR 2023
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
