Towards Balanced Alignment: Modal-Enhanced Semantic Modeling for Video Moment Retrieval
Zhihang Liu, Jun Li, Hongtao Xie, Pandeng Li, Jiannan Ge, Sun'ao Liu, Guoqing Jin
Abstract
Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since they ignore the modality imbalance problem, i.e., the semantic richness inherent in videos far exceeds that of a given limited-length sentence. Therefore, in pursuit of better alignment, a natural idea is enhancing the video modality to filter out query-irrelevant semantics, and enhancing the text modality to capture more segment-relevant knowledge. In this paper, we introduce Modal-Enhanced Semantic Modeling (MESM), a novel framework for more balanced alignment through enhancing features at two levels. First, we enhance the video modality at the frame-word level through word reconstruction. This strategy emphasizes the portions associated with query words in frame-level features while suppressing irrelevant parts. Therefore, the enhanced video contains less redundant semantics and is more balanced with the textual modality. Second, we enhance the textual modality at the segment-sentence level by learning complementary knowledge from context sentences and ground-truth segments. With the knowledge added to the query, the textual modality thus maintains more meaningful semantics and is more balanced with the video modality. By implementing two levels of MESM, the semantic information from both modalities is more balanced to align, thereby bridging the modality gap. Experiments on three widely used benchmarks, including the out-of-distribution settings, show that the proposed framework achieves a new start-of-the-art performance with notable generalization ability (e.g., 4.42% and 7.69% average gains of R1@0.7 on Charades-STA and Charades-CG). The code will be available at https://github.com/lntzm/MESM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a00aab62-d7d7-4bc8-a6c2-fb6668f11627Cited by top-tier papers15
- Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingBoqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin WangCVPR 2024 · 9 citations
- Multi-Scale Contrastive Learning for Video Temporal GroundingThong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu et al.AAAI 2025 · 7 citations
- DreamRelation: Relation-Centric Video CustomizationYujie Wei, Shiwei Zhang, Hangjie Yuan, Biao Gong et al.ICCV 2025 · 5 citations
- Generative Video Diffusion for Unseen Novel Semantic Video Moment RetrievalDezhao Luo, Shaogang Gong, Jiabo Huang, Hailin Jin et al.AAAI 2025 · 4 citations
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma et al.ICCV 2025 · 4 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-ExpertsHangbo Bao, Wenhui Wang, Li Dong, Qiang Liu et al.NeurIPS 2022 · 790 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
Related papers
- Learning Semantic Alignment with Global Modality Reconstruction for Video-Language Pre-training towards RetrievalMingchao Li, Xiaoming Shi, Haitao Leng, Wei Zhou et al.AAAI 2023 · 4 citations
- Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalDaizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan et al.ACM MM 2023 · 10 citations
- CDTR: Semantic Alignment for Video Moment Retrieval Using Concept Decomposition TransformerRan Ran, Jiwei Wei, Xiangyi Cai, Xiang Guan et al.AAAI 2025 · 6 citations
- Towards Generalisable Video Moment Retrieval: Visual-Dynamic Injection to Image-Text Pre-TrainingDezhao Luo, Jiabo Huang, Shaogang Gong, Hailin Jin et al.CVPR 2023
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
