A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction using Cluster-Based Associative Array for Selective KV Accessing
Zikang Zhou, Kaiqi Chen, Xuyang Duan, Jun Han
摘要
Attention-based LLMs excel in text generation but face redundant computations in autoregressive token generation. While KV cache mitigates this, it introduces increased memory access overhead as sequences grow. We propose Sella, a hardware-software co-design using cluster-based associative arrays to predict Q-K correlations, enabling selective KV cache access and reducing memory access without retraining. Sella includes a specialized accelerator featuring a prediction engine to improve performance and energy efficiency. Experiments show Sella achieves , , , and speedup over SpAtten, Sanger, TITAN RTX GPU, and Xeon CPU, respectively, reducing off-chip memory access by up to with negligible accuracy loss. -Large Language Models, KV Cache, Accelerator
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 被引用 6 次
- KVO-LLM: Boosting Long-Context Generation Throughput for Batched LLM InferenceZhenyu Li, Dongxu Lyu, Gang Wang, Yuzhou Chen 等DAC 2025 · 被引用 1 次
- LAD: Efficient Accelerator for Generative Inference of LLM with Locality Aware DecodingHaoran Wang, Yuming Li, Haobo Xu, Ying Wang 等HPCA 2025 · 被引用 6 次
- Reducing Transformer Key-Value Cache Size with Cross-Layer AttentionWilliam Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda 等NeurIPS 2024 · 被引用 150 次
- ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV CachingYoupeng Zhao, Di Wu, Jun WangISCA 2024 · 被引用 35 次
