Revealing Single Frame Bias for Video-and-Language Learning
Jie Lei, Tamara L. Berg, Mohit Bansal
Abstract
Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong “static appearance bias” in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46487631-7eef-4794-8bad-2fb3b949c1dbCited by top-tier papers65
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
- VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetSihan Chen, Handong Li, Qunbo Wang, Zijia Zhao et al.NeurIPS 2023 · 246 citations
- HiTeA: Hierarchical Temporal-Aware Video-Language Pre-trainingQinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu et al.ICCV 2023 · 102 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- RTQ: Rethinking Video-language Understanding Based on Image-text ModelXiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang et al.ACM MM 2023 · 14 citations
- SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-trainingYuanze Lin, Chen Wei, Huiyu Wang, Alan L. Yuille et al.ICCV 2023 · 18 citations
- Revisiting the "Video" in Video-Language UnderstandingShyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu et al.CVPR 2022 · 121 citations
- STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-trainingWeihong Zhong, Mao Zheng, Duyu Tang, Xuan Luo et al.AAAI 2023 · 9 citations
- LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal ModelingDongsheng Chen, Chaofan Tao, Lu Hou, Lifeng Shang et al.EMNLP 2022 · 11 citations
