VRoPE: Rotary Position Embedding for Video Large Language Models
Zikang Liu, Longteng Guo, Yepeng Tang, Tongtian Yue, Junxian Cai, Kai Ma, Qingbin Liu, Xi Chen, Jing Liu
Abstract
Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames. Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensions separately but suffer from two major limitations: positional bias in attention distribution and disruptions in video-text transitions. To overcome these issues, we propose Video Rotary Position Embedding (VRoPE), a novel positional encoding method tailored for Video-LLMs. Specifically, we introduce a more balanced encoding strategy that mitigates attention biases, ensuring a more uniform distribution of spatial focus. Additionally, our approach restructures positional indices to ensure a smooth transition between video and text tokens. Extensive experiments on different models demonstrate that VRoPE consistently outperforms previous RoPE variants, achieving significant improvements in video understanding, temporal reasoning, and retrieval tasks. Code is available at https://github.com/johncaged/VRoPE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68b01e4f-fdc6-439e-8276-c708acd1432bCited by top-tier papers12
- Revisiting Multimodal Positional Encoding in Vision–Language ModelsJie Huang, Xuejing Liu, Sibo Song, RuiBing Hou et al.ICLR 2026 · 20 citations
- RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion TransformersAhmet Berke Gökmen, Yigit Ekin, Bahri Batuhan Bilecen, Aysegul DundarNeurIPS 2025 · 14 citations
- A Circular Argument: Does RoPE need to be Equivariant for Vision?Chase van de Geijn, Timo Lüddecke, Polina Turishcheva, Alexander S. EckerNeurIPS 2025 · 6 citations
- VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information AssumptionTianxiong Zhong, Xingye Tian, Boyuan Jiang, Xuebo Wang et al.NeurIPS 2025 · 4 citations
- Causality Matters: How Temporal Information Emerges in Video Language ModelsYumeng Shi, Quanyu Long, Yin Wu, Wenya WangAAAI 2026 · 4 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 279 citations
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsSoravit Changpinyo, Piyush Sharma, Nan Ding, Radu SoricutCVPR 2021
- Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video UnderstandingPeng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao et al.CVPR 2024
Related papers
- TC-LLaVA: Rethinking the Transfer of LLava from Image to Video Understanding with Temporal ConsiderationsMingze Gao, Jingyu Liu, Mingda Li, Jiangtao Xie et al.AAAI 2025 · 4 citations
- VideoRoPE: What Makes for Good Video Rotary Position Embedding?Xilin Wei, Xiaoran Liu, Yuhang Zang, Xiaoyi Dong et al.ICML 2025
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language ModelsChengcheng Wang, Jianyuan Guo, Hongguang Li, Yuchuan Tian et al.ICML 2026 · 14 citations
- 3D-RPE: Enhancing Long-Context Modeling Through 3D Rotary Position EncodingXindian Ma, Wenyuan Liu, Peng Zhang, Nan XuAAAI 2025 · 14 citations
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang et al.ICML 2026
