Learning Beyond Still Frames: Scaling Vision-Language Models with Video
Yiyuan Zhang, Handong Jing, Jing Liu, Xiangyu Yue
Abstract
High-quality image-text data is critical for Vision-Language Models (VLMs), yet traditional image-based pretraining is resource-intensive and fails to capture the temporal dynamics needed for video understanding. To address this, we introduce video pretraining to enhance VLMs with temporal reasoning. We propose Causal Hierarchical Aggregation, a novel method that efficiently processes video by separating computationally heavy spatial encoding from lightweight temporal propagation. This technique builds hierarchical receptive fields, enabling effective learning from large-scale video data. Scaling our method to over 100 billion video tokens, we achieve state-of-the-art performance and high throughput on both image and video understanding tasks (Figure 1). Our approach offers a scalable solution to advance multimodal learning for dynamic contexts. Our code and pretrained models will be released at https://github.com/invictus717/LLaVA-Prime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e496677-258c-48aa-8a7b-b48467b4f0abCited by top-tier papers1
Ask how each one uses itBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu et al.ICML 2024 · 94 citations
- LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal ModelingDongsheng Chen, Chaofan Tao, Lu Hou, Lifeng Shang et al.EMNLP 2022 · 11 citations
- HiVLP: Hierarchical Interactive Video-Language Pre-TrainingBin Shao, Jianzhuang Liu, Renjing Pei, Songcen Xu et al.ICCV 2023 · 6 citations
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question DecompositionYiyang Huang, Yizhou Wang, Yun FuEMNLP 2025
- VidLA: Video-Language Alignment at ScaleMamshad Nayeem Rizve, Fan Fei, Jayakrishnan Unnikrishnan, Son Tran et al.CVPR 2024 · 3 citations
