Lune

ICCV2025顶会

Learning Beyond Still Frames: Scaling Vision-Language Models with Video

Yiyuan Zhang, Handong Jing, Jing Liu, Xiangyu Yue

2025年份
2被引次数
1顶会引用

摘要

High-quality image-text data is critical for Vision-Language Models (VLMs), yet traditional image-based pretraining is resource-intensive and fails to capture the temporal dynamics needed for video understanding. To address this, we introduce video pretraining to enhance VLMs with temporal reasoning. We propose Causal Hierarchical Aggregation, a novel method that efficiently processes video by separating computationally heavy spatial encoding from lightweight temporal propagation. This technique builds hierarchical receptive fields, enabling effective learning from large-scale video data. Scaling our method to over 100 billion video tokens, we achieve state-of-the-art performance and high throughput on both image and video understanding tasks (Figure 1). Our approach offers a scalable solution to advance multimodal learning for dynamic contexts. Our code and pretrained models will be released at https://github.com/invictus717/LLaVA-Prime.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖