HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
Haoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu, Ruiwen Xu
Abstract
Vision-Language Models (VLMs) have made significant progress in multimodal tasks. However, their performance often deteriorates in long-context scenarios, particularly long videos. While Rotary Position Embedding (RoPE) has been widely adopted for length generalization in Large Language Models (LLMs), extending vanilla RoPE to capture the intricate spatial-temporal dependencies in videos remains an unsolved challenge. Existing methods typically allocate different frequencies within RoPE to encode 3D positional information. However, these allocation strategies mainly rely on heuristics, lacking in-depth theoretical analysis. In this paper, we first study how different allocation strategies impact the long-context capabilities of VLMs. Our analysis reveals that current multimodal RoPEs fail to reliably capture semantic similarities over extended contexts. To address this issue, we propose HoPE, a Hybrid of Position Embedding designed to improve the long-context capabilities of VLMs. HoPE introduces a hybrid frequency allocation strategy for reliable semantic modeling over arbitrarily long contexts, and a dynamic temporal scaling mechanism to facilitate robust learning and flexible inference across diverse context lengths. Extensive experiments across four video benchmarks on long video understanding and retrieval tasks demonstrate that HoPE consistently outperforms existing methods, confirming its effectiveness. Our code is available at https://github.com/hrlics/HoPE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal GroundingChaohong Guo, Yihan He, Yongwei Nie, Fei Ma et al.CVPR 2026 · 2 citations
- Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience StrategiesPeng Gao, Yujian Lee, Xiaofeng Zhang, Zailong Chen et al.AAAI 2026
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Shunted Self-Attention via Multi-Scale Token AggregationSucheng Ren, Daquan Zhou, Shengfeng He, Jiashi Feng et al.CVPR 2022 · 326 citations
Related papers
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan et al.CVPR 2026 · 3 citations
- HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationYuhan Chen, Ang Lv, Jian Luan, Bin Wang et al.ACL 2025
- Revisiting Multimodal Positional Encoding in Vision–Language ModelsJie Huang, Xuejing Liu, Sibo Song, RuiBing Hou et al.ICLR 2026 · 20 citations
- VRoPE: Rotary Position Embedding for Video Large Language ModelsZikang Liu, Longteng Guo, Yepeng Tang, Tongtian Yue et al.EMNLP 2025 · 1 citation
- VideoRoPE: What Makes for Good Video Rotary Position Embedding?Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong et al.ICML 2025
