Lune

ICCV2025Top-tier venue

Learning Beyond Still Frames: Scaling Vision-Language Models with Video

Yiyuan Zhang, Handong Jing, Jing Liu, Xiangyu Yue

2025Year
2Citations
1Top-tier citations

Abstract

High-quality image-text data is critical for Vision-Language Models (VLMs), yet traditional image-based pretraining is resource-intensive and fails to capture the temporal dynamics needed for video understanding. To address this, we introduce video pretraining to enhance VLMs with temporal reasoning. We propose Causal Hierarchical Aggregation, a novel method that efficiently processes video by separating computationally heavy spatial encoding from lightweight temporal propagation. This technique builds hierarchical receptive fields, enabling effective learning from large-scale video data. Scaling our method to over 100 billion video tokens, we achieve state-of-the-art performance and high throughput on both image and video understanding tasks (Figure 1). Our approach offers a scalable solution to advance multimodal learning for dynamic contexts. Our code and pretrained models will be released at https://github.com/invictus717/LLaVA-Prime.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4e496677-258c-48aa-8a7b-b48467b4f0ab

Cited by top-tier papers1

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines