Clover: Towards A Unified Video-Language Alignment and Fusion Model
Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, Rongrong Ji
Abstract
Building a universal Video-Language model for solving various video understanding tasks (e.g., text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal, most recent works build the model by stacking uni-modal and cross-modal feature encoders and train it with pair-wise contrastive pre-text tasks. Though offering attractive generality, the resulted models have to compromise between efficiency and performance. They mostly adopt different architectures to deal with different downstream tasks. We find this is because the pair-wise training cannot well align and fuse features from different modalities. We then introduce Clover-a Correlated Video-Language pre-training method-towards a universal Video-Language model for solving multiple video understanding tasks with neither performance nor efficiency compromise. It improves cross-modal feature alignment and fusion via a novel tri-modal alignment pre-training task. Additionally, we propose to enhance the tri-modal alignment via incorporating learning from semantic masked samples and a new pair-wise ranking loss. Clover establishes new state-of-the-arts on multiple downstream tasks, including three retrieval tasks for both zero-shot and finetuning settings, and eight video question answering tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd97ce49-19a3-42dd-9bde-065dc8f89c4eCited by top-tier papers17
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalJiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu et al.CVPR 2024 · 52 citations
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn et al.CVPR 2026 · 33 citations
- Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited ModalitiesAdriel Saporta, Aahlad Manas Puli, Mark Goldstein, Rajesh RanganathNeurIPS 2024 · 28 citations
- Diffusion-Inspired Truncated Sampler for Text-Video RetrievalJiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan et al.NeurIPS 2024 · 16 citations
Builds on25
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
Related papers
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- LAVENDER: Unifying Video-Language Understanding as Masked Language ModelingLinjie Li, Zhe Gan, Kevin Lin, Chung-Ching Lin et al.CVPR 2023
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi et al.ICML 2023 · 237 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- Align and Prompt: Video-and-Language Pre-training with Entity PromptsDongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles et al.CVPR 2022
