IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
Yi Zhao, Zuchao Li, Hai Zhao
Abstract
LLMs encounter significant challenges in resource consumption nowadays, especially with long contexts. Despite extensive efforts dedicate to enhancing inference efficiency, these methods primarily exploit internal sparsity within the models, without leveraging external information for optimization. We identify the high similarity of attention matrices across different-scale LLMs, which offers a novel perspective for optimization. We first conduct a comprehensive analysis of how to measure similarity, how to select mapping Layers and whether mapping is consistency. Based on these insights, we introduce the IAM framework, which achieves dual benefits of accelerated attention computation and reduced KV cache usage by performing attention mapping between small and large LLMs. Our experimental results demonstrate that IAM can accelerate prefill by 15% and reduce KV cache usage by 22.1% without appreciably sacrificing performance. Experiments on different series of models show the generalizability of IAM. Importantly, it is also orthogonal to many existing KV cache optimization methods, making it a versatile addition to the current toolkit for enhancing LLM efficiency. Our code is available at https://github.com/QQQ-yi/IAM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9dfd3e79-8158-4b48-bfb3-7bd50911d789Cited by top-tier papers2
- From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic HorizonsXiangyu Ma, Teng Xiao, Zuchao Li, Lefei ZhangACL 2026
- From Parameters to Performance: A Data-Driven Study on LLM Structure and DevelopmentSuqing Wang, Zuchao Li, Luohe Shi, Bo Du et al.EMNLP 2025
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Q Cache: Visual Attention Is Valuable in Less than Half of Decode Layers for Multimodal Large Language ModelJiedong Zhuang, Lu Lu, Ming Dai, Rui Hu et al.AAAI 2026
- SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep LayersZicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi et al.ACL 2025 · 7 citations
- SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal BudgetZihao Wang, Bin Cui, Shaoduo GanICLR 2025
- ZipVL: Accelerating Vision-Language Models Through Dynamic Token SparsityYefei He, Feng Chen, Jing Liu, Wenqi Shao et al.ICCV 2025 · 1 citation
- Layer-Condensed KV Cache for Efficient Inference of Large Language ModelsHaoyi Wu, Kewei TuACL 2024
