Align and Prompt: Video-and-Language Pre-training with Entity Prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, Steven C. H. Hoi
Abstract
Video-and-language pre-training has shown promising improvements on various downstream tasks. Most previous methods capture cross-modal interactions with a standard transformer-based multimodal encoder, not fully addressing the misalignment between unimodal video and text features. Besides, learning fine-grained visual-language alignment usually requires off-the-shelf object detectors to provide object information, which is bottlenecked by the detector's limited vocabulary and expensive computation cost.
In this paper, we propose Align and Prompt: a new video-and-language pre-training framework (ALPRO), which operates on sparsely-sampled video frames and achieves more effective cross-modal alignment without explicit object detectors. First, we introduce a video-text contrastive (VTC) loss to align unimodal video-text features at the instance level, which eases the modeling of cross-modal interactions. Then, we propose a novel visually-grounded pre-training task, prompting entity modeling (PEM), which learns fine-grained alignment between visual region and text entity via an entity prompter module in a self-supervised way. Finally, we pretrain the video-and-language transformer models on large webly-source video-text pairs using the proposed VTC and PEM losses as well as two standard losses of masked language modeling (MLM) and video-text matching (VTM). The resulting pre-trained model achieves state-of-the-art performance on both text-video retrieval and videoQA, outperforming prior work by a substantial margin. Implementation and pre-trained models are available at https://github.com/salesforce/ALPRO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbfa7ac2-869b-455c-adc8-5b8940a3add3Cited by top-tier papers78
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi et al.ICML 2023 · 237 citations
- OmniVL: One Foundation Model for Image-Language and Video-Language TasksJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo et al.NeurIPS 2022 · 205 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu et al.ICCV 2023 · 6 citations
- Structured Video-Language Modeling with Temporal Grouping and Spatial GroundingYuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang et al.ICLR 2024
- PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language ModelsYuan Yao, Qianyu Chen, Ao Zhang, Wei Ji et al.EMNLP 2022 · 33 citations
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi et al.ACL 2021
- Object-aware Video-language Pre-training for RetrievalAlex Jinpeng Wang, Yixiao Ge, Guanyu Cai, Rui Yan et al.CVPR 2022 · 65 citations
