VideoPrism: A Foundational Visual Encoder for Video Understanding
Long Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J. Sun, Luke Friedman, Rui Qian, Tobias Weyand, Yue Zhao, Rachel Hornung, Florian Schroff
Abstract
We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M video clips with noisy parallel text (e.g., ASR transcripts). The pretraining approach improves upon masked autoencoding by globallocal distillation of semantic video embeddings and a token shuffling scheme, enabling Video-Prism to focus primarily on the video modality while leveraging the invaluable text associated with videos. We extensively test VideoPrism on four broad groups of video understanding tasks, from web video question answering to CV for science, achieving state-of-the-art performance on 31 out of 33 video understanding benchmarks. Our models are released at https://github. com/google-deepmind/videoprism . * Equal primary contribution. † Equal core technical contribution. ‡ Equal senior contribution, project leads. § This work was done while the author was a student researcher at Google Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94d7eb8a-e0e1-4071-838c-1f0c22177629Cited by top-tier papers48
- GenRL: Multimodal-foundation world models for generalization in embodied agentsPietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C. Courville et al.NeurIPS 2024 · 37 citations
- AToken: A Unified Tokenizer for VisionJiasen Lu, Liangchen Song, Mingze Xu, Byeongjoo Ahn et al.CVPR 2026 · 33 citations
- VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingYongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng et al.AAAI 2025 · 27 citations
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence LearningApoorv Vyas, Heng-Jui Chang, Cheng-Fu Yang, Po-Yao Huang et al.CVPR 2026 · 23 citations
- TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-AlignmentWei Li, Hehe Fan, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2024 · 18 citations
Builds on64
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He et al.CVPR 2024 · 18 citations
- End-to-end Generative Pretraining for Multimodal Video CaptioningPaul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia SchmidCVPR 2022 · 152 citations
- Structured Video-Language Modeling with Temporal Grouping and Spatial GroundingYuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang et al.ICLR 2024
- Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsHongwei Xue, Tiankai Hang, Yanhong Zeng, Yuchong Sun et al.CVPR 2022 · 107 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
