Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Vision-Language Models
Chengcheng Wang, Jianyuan Guo, Hongguang Li, Yuchuan Tian, Ying Nie, Chang Xu, Kai Han
摘要
Rotary Position Embedding (RoPE) is widely adopted in large language models, but when applied to vision-language models (VLMs) it couples text and image position indices and can introduce spurious cross-modal relative-position bias. We propose Per-Token Distance (PTD) to quantify cross-modal positional disentanglement, and we prove that is a sufficient condition to eliminate the geometric attention bias induced by RoPE. Guided by this criterion, we introduce Circle-RoPE, which remaps 2D image-token coordinates onto an annulus orthogonal to the text position axis, yielding a cone-like geometry where each text token is equidistant to all image tokens while preserving intra-image spatial structure. We further propose Alternating Geometry Encoding (AGE) to synergize complementary geometric priors by alternating the decoupled geometry of Circle-RoPE and the grid-based prior of standard RoPE across layers. This design ensures both rigorous cross-modal disentanglement and the preservation of fine-grained intra-image spatial structure, and experiments on diverse VLM backbones and multimodal benchmarks show consistent gains in spatial grounding and visual reasoning. The code is available at https://github.com/lose4578/CircleRoPE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Revisiting Multimodal Positional Encoding in Vision–Language ModelsJie Huang, Xuejing Liu, Sibo Song, RuiBing Hou 等ICLR 2026 · 被引用 20 次
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan 等CVPR 2026 · 被引用 3 次
- MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language ModelsRuoxiang Huang, Zhen YuanCVPR 2026 · 被引用 2 次
它引用的顶会 Paper11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 等ACL 2025 · 被引用 377 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
- Eve: Efficient Multimodal Vision Language Models with Elastic Visual ExpertsMiao Rang, Zhenni Bi, Chuanjian Liu, Yehui Tang 等AAAI 2025 · 被引用 16 次
- VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic ModelWenhao Li, Xiu Su, Yichao Cao, Hongyan Xu 等ICML 2026 · 被引用 13 次
相关 Paper
- VRoPE: Rotary Position Embedding for Video Large Language ModelsZikang Liu, Longteng Guo, Yepeng Tang, Tongtian Yue 等EMNLP 2025 · 被引用 1 次
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D PlaneHaoyu Liu, Sucheng Ren, Tingyu Zhu, Peng Wang 等ICML 2026
- Mitigating Object Hallucination via Concentric Causal AttentionYun Xing, Yiheng Li, Ivan Laptev, Shijian LuNeurIPS 2024 · 被引用 78 次
- An Anchor-based Relative Position Embedding Method for Cross-Modal TasksYa Wang, Xingwu Sun, Fengzong Lian, Zhanhui Kang 等EMNLP 2022 · 被引用 1 次
- IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token PruningZhichao Sun, Yidong Ma, Gang Liu, Nemo Chen 等ICLR 2026 · 被引用 11 次
