V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
Junqi Ge, Ziyi Chen, Jintao Lin, Jinguo Zhu, Xihui Liu, Jifeng Dai, Xizhou Zhu
摘要
Vision-Language Models (VLMs) have shown promising capabilities in handling various multimodal tasks, yet they struggle in long-context scenarios, particularly in tasks involving videos, high-resolution images, or lengthy imagetext documents. In our work, we first conduct an empirical analysis of the long-context capabilities of VLMs using our augmented long-context multimodal datasets. Our findings reveal that directly applying the positional encoding mechanism used for textual tokens to visual tokens is suboptimal, and VLM performance degrades sharply when the position encoding exceeds the model's context window. To address this, we propose Variable Visual Position Encoding (V2PE), a novel positional encoding approach that employs variable and smaller increments for visual tokens, enabling more efficient management of long multimodal sequences. Our experiments demonstrate the effectiveness of V2PE to enhances VLMs' ability to effectively understand and reason over long multimodal contexts. We further integrate V2PE with our augmented long-context multimodal datasets to fine-tune the open-source VLM, InternVL2. The fine-tuned model achieves strong performance on both standard and long-context multimodal tasks. Notably, when the sequence length of the training dataset is increased to 256K tokens, the model is capable of processing multimodal sequences up to 1M tokens, highlighting its potential for realworld long-context applications. The code and models will be available at https://github.com/OpenGVLab/V2PE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Revisiting Multimodal Positional Encoding in Vision–Language ModelsJie Huang, Xuejing Liu, Sibo Song, RuiBing Hou 等ICLR 2026 · 被引用 20 次
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit MapsSicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong 等CVPR 2026 · 被引用 19 次
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless LoopsJiyuan Fu, Kaixun Jiang, Lingyi Hong, Jinglun Li 等ICLR 2026 · 被引用 12 次
- LoVR: A Benchmark for Long Video Retrieval in Multimodal ContextsHao Liang, Qifeng Cai, Zhaoyang Han, Hejun Dong 等WWW 2026 · 被引用 11 次
- Self-alignment of Large Video Language Models with Refined Regularized Preference OptimizationPritam Sarkar, Ali EtemadNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper53
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Design Choices for Extending the Context Length of Visual Language ModelsMukai Li, Lei Li, Shansan Gong, Qi LiuACL 2025
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language ModelsHaoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu 等NeurIPS 2025 · 被引用 4 次
- TULIP: Token-length Upgraded CLIPIvona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki M. Asano, Nanne van Noord 等ICLR 2025
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang 等NeurIPS 2025 · 被引用 69 次
- PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series in Typhoon ForecastingHaobo Li, Eunseo Jung, Zixin Chen, Zhaowei Wang 等NeurIPS 2025
