InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
Weilin Zhao, Zihan Zhou, Zhou Su, Chaojun Xiao, Yuxuan Li, Yanghao Li, Yudi Zhang, Weilun Zhao, Zhen Li, Yuxiang Huang, Ao Sun, Xu Han, Zhiyuan Liu
摘要
Long-sequence processing is a critical capability for modern large language models. However, the self-attention mechanism in the standard Transformer architecture faces severe computational and memory bottlenecks when processing long sequences. While trainable sparse attention methods offer a promising solution, existing approaches such as NSA introduce excessive extra parameters and disrupt the conventional pretrain-on-short, finetune-on-long workflow, resulting in slow convergence and difficulty in acceleration. To overcome these limitations, we introduce dense-sparse switchable attention framework, termed as InfLLM-V2. InfLLM-V2 is a trainable sparse attention that seamlessly adapts models from short to long sequences. Specifically, InfLLM-V2 reuses dense attention parameters through parameter-free architecture modification, maintaining consistency between short and long sequence processing. Additionally, InfLLM-V2 ensures computational efficiency across all sequence lengths, by using dense attention for short inputs and smoothly transitioning to sparse attention for long sequences. To achieve practical acceleration, we further introduce an efficient implementation of InfLLM-V2 that significantly reduces the computational overhead. Our experiments on long-context understanding and chain-of-thought reasoning demonstrate that InfLLM-V2 is 4 faster than dense attention while retaining 98.1% and 99.7% of the performance, respectively. Based on the InfLLM-V2 framework, we have trained and open-sourced MiniCPM4.1 (https://huggingface.co/openbmb/MiniCPM4.1-8B), a hybrid reasoning model, providing a reproducible implementation for the research community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient TransformersZecheng Tang, Quantong Qiu, Yi Yang, Zhiyi Hong 等ICML 2026 · 被引用 4 次
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature SpaceZhenyi Shen, Junru Lu, Lin Gui, Jiazheng Li 等ICML 2026 · 被引用 2 次
- Learning When to Attend: Conditional Memory Access for Long-Context LLMsSakshi Choudhary, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez 等ICML 2026 · 被引用 2 次
- ECHO: Efficient KV Cache Offloading with Lossless Prefetching for Serving Native Sparse Attention LLMsGuangda Liu, Wenhao Chen, Chengwei Li, Zhenyu Ning 等OSDI 2026
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- A Unified Sparse Attention via Multi-Granularity CompressionSiran Liu, Zheng Cao, Yongchao HeICML 2026
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 等ACL 2025 · 被引用 334 次
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng 等DAC 2022 · 被引用 37 次
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence InferenceXunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma 等ICLR 2025
