UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM Inference
Weikai Xu, Wenxuan Zeng, Qianqian Huang, Meng Li, Ru Huang
Abstract
Transformer-based large language models (LLMs) have achieved impressive performance in various natural language processing (NLP) applications. However, the high memory and computation cost induced by the KV cache limits the inference efficiency, especially for long input sequences. Compute-in-memory (CIM)-based accelerators have been proposed for LLM acceleration with KV cache pruning. However, as existing accelerators only support static pruning with a fixed pattern or dynamic pruning with primitive implementations, they suffer from either high accuracy degradation or low efficiency. In this paper, we propose a ferroelectric FET (FeFET)-based unified content addressable memory (CAM) and CIM architecture, dubbed as UniCAIM. UniCAIM features simultaneous support for static and dynamic pruning with 3 computation modes: 1) in the CAM mode, UniCAIM enables approximate similarity measurement in time for dynamic KV cache pruning with high energy efficiency; 2) in the charge-domain CIM mode, static pruning can be supported based on accumulative similarity score, which is much more flexible compared to fixed patterns; 3) in the current-domain mode, exact attention computation can be conducted with a subset of selected KV cache. We further propose a novel CAM/CIM cell design that leverages the multi-level characteristics of FeFETs for signed multibit storage of the KV cache and in-place attention computation. With extensive experimental results, we demonstrate UniCAIM can reduce the area-energy-delay product (AEDP) by over the state-of-the-art CIM-based LLM accelerators at the circuit level, along with high accuracy comparable with dense attention at the application level, showing its great potential for efficient long-context LLM inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f3d8718-379f-4071-ad65-9353c8377eabBuilds on15
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
Related papers
- VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-MemoryXuchu Huang, Haonan Du, Min Zhou, Zheyu Yan et al.DAC 2025 · 3 citations
- Compact and Efficient CAM Architecture through Combinatorial Encoding and Self-Terminating Searching for In-Memory-Searching AcceleratorWeikai Xu, Jin Luo, Qianqian Huang, Ru HuangDAC 2024 · 1 citation
- Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memoryJiahao Cai, Mohsen Imani, Kai Ni, Grace Li Zhang et al.DAC 2022 · 17 citations
- Energy Efficient Dual Designs of FeFET-Based Analog In-Memory Computing with Inherent Shift-Add CapabilityZeyu Yang, Qingrong Huang, Yu Qian, Kai Ni et al.DAC 2024 · 9 citations
- FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory ComputingChao Li, Zhicheng Xu, Bo Wen, Ruibin Mao et al.DAC 2024 · 6 citations
