Rethinking Video-Language Model from the Language Input Perspective
Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu, Daizong Liu
Abstract
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefined by the specific template. In real-world applications, such a strict assumption is impossible to satisfy since 1) predefining all the texts is extremely time-consuming and labor-intensive. 2) these predefined text inputs are too restrictive and user-unfriendly, limiting their applications. It is observed that given a video input, texts with similar semantics but different templates lead to various performances. To this end, in this paper, we propose a novel plug-and-play framework for various VLM-based methods to fully bridge videos and texts. Specifically, we first generate positive and negative texts from the original ones to target specific text components. Then, we propose an attribute-based text reasoning strategy to mine fine-grained textual semantics of generated texts. Finally, we utilize videos as guidance to conduct cross-modal bridging by designing a self-weighted loss. Extensive experiments show that the proposed method can serve as the plug-and-play module to effectively improve the performance of state-of-the-art VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4510cd78-508b-4a0f-add1-47cc7346027bCited by top-tier papers8
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo et al.ICLR 2026 · 45 citations
- Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World TrustworthinessXiang Fang, Wanlong Fang, Wei JiICML 2026 · 17 citations
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 17 citations
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using LanguageXiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu et al.ACM MM 2024 · 8 citations
- CARE: Covariance-Aware and Rank-Enhanced Decomposition for Enabling Multi-Head Latent AttentionZhongzhu Zhou, Fengxiang Bie, Ziyan Chen, Zhenyu Zhang et al.ICLR 2026 · 4 citations
Builds on26
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang et al.AAAI 2020 · 170 citations
- Negative Sample Matters: A Renaissance of Metric Learning for Temporal GroundingZhenzhi Wang, Limin Wang, Tao Wu, Tianhao Li et al.AAAI 2022 · 170 citations
- Learn from Relational Correlations and Periodic Events for Temporal Knowledge Graph ReasoningKe Liang, Lingyuan Meng, Meng Liu, Yue Liu et al.SIGIR 2023 · 117 citations
- COSTA: Covariance-Preserving Feature Augmentation for Graph Contrastive LearningYifei Zhang, Hao Zhu, Zixing Song, Piotr Koniusz et al.KDD 2022 · 95 citations
Related papers
- Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language ModelsWenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang et al.CVPR 2023
- PlanLLM: Video Procedure Planning with Refinable Large Language ModelsDejie Yang, Zijing Zhao, Yang LiuAAAI 2025 · 8 citations
- Text-Adaptive Multiple Visual Prototype Matching for Video-Text RetrievalChengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang et al.NeurIPS 2022 · 52 citations
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar et al.AAAI 2025 · 3 citations
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 53 citations
