Towards Long-window Anchoring in Vision-Language Model Distillation
Haoyi Zhou, Shuo Li, Tianyu Chen, Qi Song, Chonghan Gao, Jianxin Li
Abstract
While large vision-language models (VLMs) demonstrate impressive long-context understanding, their prevalent small branches fails on linguistics-photography alignment for limited window size. We discover that knowledge distillation improve students capability as compelementary to Rotary Position Embeddings (RoPE) on certain windows size (anchored from large models). Building on this insight, we propose LAid, which explicitly targets the transfer of long-range attention mechanisms through two complementary components: (1) a progressive distance-weighted attention matching that dynamically emphasizes longer position differences during training, and (2) a learnable RoPE response gain modulation that selectively amplifies position sensitivity where needed. Extensive experiments across multiple model families demonstrate that LAid-distilled models achieve up to 3.2× longer effective context windows compared to baseline small models, while maintaining or improving performance on standard VL benchmarks. Spectral analysis also suggests that LAid successfully preserves crucial low-frequency attention components that conventional methods fail to transfer. Our work not only provides practical techniques for building more efficient long-context VLMs but also offers theoretical insights into how positional understanding emerges and transfers during distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15dcbb1e-bf7f-4e6c-b309-1b6594bea048Builds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- LongRoPE: Extending LLM Context Window Beyond 2 Million TokensYiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu et al.ICML 2024 · 316 citations
Related papers
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language ModelsHaoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu et al.NeurIPS 2025 · 4 citations
- Revisiting Multimodal Positional Encoding in Vision–Language ModelsJie Huang, Xuejing Liu, Sibo Song, RuiBing Hou et al.ICLR 2026 · 20 citations
- Wavelet-based Positional Representation for Long ContextYui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko SaitoICLR 2025
- Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free MethodXinhao Xu, Jiaxin Li, Hui Chen, Zijia Lin et al.ACL 2025 · 1 citation
- No Head Left Behind - Multi-Head Alignment Distillation for TransformersTianyang Zhao, Kunwar Yashraj Singh, Srikar Appalaraju, Peng Tang et al.AAAI 2024 · 5 citations
