Design Choices for Extending the Context Length of Visual Language Models
Mukai Li, Lei Li, Shansan Gong, Qi Liu
摘要
Visual Language Models (VLMs) demonstrate impressive capabilities in processing multimodal inputs, yet applications such as visual agents, which require handling multiple images and high-resolution videos, demand enhanced long-range modeling. Moreover, existing opensource VLMs lack systematic exploration into extending their context length, and commercial models often provide limited details. To tackle this, we aim to establish an effective solution that enhances long context performance of VLMs while preserving their capacities in short context scenarios. Towards this goal, we make the best design choice through extensive experiment settings from data curation to context window extending and utilizing: ( 1) we analyze data sources and length distributions to construct ETVLM -a data recipe to balance the performance across scenarios; (2) we examine existing position extending methods, identify their limitations and propose M-RoPE++ as an enhanced approach; we also choose to solely instruction-tune the backbone with mixed-source data; (3) we discuss how to better utilize extended context windows and propose hybrid-resolution training. Built on the Qwen-VL series model, we propose GI-RAFFE, which is effectively extended to 128K lengths. Evaluated on extensive long context VLM benchmarks such as VideoMME and Viusal Haystacks, our GIRAFFE achieves stateof-the-art performance among similarly sized open-source long VLMs and is competitive with commercial model GPT-4V. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position EncodingJunqi Ge, Ziyi Chen, Jintao Lin, Jinguo Zhu 等ICCV 2025 · 被引用 5 次
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language ModelsHaoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu 等NeurIPS 2025 · 被引用 4 次
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang 等NeurIPS 2025 · 被引用 69 次
- Glyph: Scaling Context Windows via Visual-Text CompressionJiale Cheng, Yusen Liu, Xinyu Zhang, Yulin Fei 等ACL 2026
- Visual Context Window Extension: A New Perspective for Long Video UnderstandingHongchen Wei, Zhenzhong ChenACM MM 2025
