Lune

DAC2025顶会

A Memory-Efficient LLM Accelerator with Q-K Correlation Prediction using Cluster-Based Associative Array for Selective KV Accessing

Zikang Zhou, Kaiqi Chen, Xuyang Duan, Jun Han

2025年份
2被引次数

摘要

Attention-based LLMs excel in text generation but face redundant computations in autoregressive token generation. While KV cache mitigates this, it introduces increased memory access overhead as sequences grow. We propose Sella, a hardware-software co-design using cluster-based associative arrays to predict Q-K correlations, enabling selective KV cache access and reducing memory access without retraining. Sella includes a specialized accelerator featuring a prediction engine to improve performance and energy efficiency. Experiments show Sella achieves 2.1×2.1 \times, 93.8×93.8 \times, 31.4×31.4 \times, and 53.5×53.5 \times speedup over SpAtten, Sanger, TITAN RTX GPU, and Xeon CPU, respectively, reducing off-chip memory access by up to 66%66 \% with negligible accuracy loss. -Large Language Models, KV Cache, Accelerator

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get 4f3a1d15-1a62-42dc-9403-5d59210c499c

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖