VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-Memory
Xuchu Huang, Haonan Du, Min Zhou, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
Abstract
Transformer models have achieved state-of-the-art performance in various natural language processing (NLP) and computer vision (CV) tasks. To meet their substantial computational demands, the compute-in-memory (CiM) architectures, which alleviate the memory wall problem and enable efficient vector-matrix multiplication (VMM), have been adopted for transformer accelerators. However, the dynamic VMM involved in the attention mechanism, which necessitates runtime write operations, presents significant challenges for non-volatile memory (NVM)-based CiM designs. High write overhead, complex compute-write-compute (CWC) dependencies, and limited endurance reduce their effectiveness. In this paper, we propose VQT-CiM, a ferroelectric FET (FeFET)-based CiM design that accelerates vector quantization (VQ) enhanced transformers by eliminating the runtime write operations. VQT-CiM quantizes keys and values in self-attention to convert dynamic VMMs in inner-product and weighted-sum into static VMMs with the codebooks, enabling efficient calculations with CiM crossbars. However, directly applying VQ hinders the accuracy of transformer model due to its limited representation capability. To address this, we introduce a vector quantization scheme that integrates residual VQ (RVQ) and product VQ (PVQ) for enhanced representation space. We present an efficient hardware implementation for the proposed VQT-CiM with optimized dataflow in RVQ, which incorporates the FeFET-based CiM crossbars and peripheral digital circuits. Evaluation results suggest that VQT-CiM achieves the and improvements in energy efficiency and throughput, respectively, compared to state-of-the-art NVM-based CiM transformer designs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 02add2a0-06fd-425f-be16-097eb30d07b1Cited by top-tier papers1
Ask how each one uses itRelated papers
- UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM InferenceWeikai Xu, Wenxuan Zeng, Qianqian Huang, Meng Li et al.DAC 2025 · 3 citations
- Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memoryJiahao Cai, Mohsen Imani, Kai Ni, Grace Li Zhang et al.DAC 2022 · 17 citations
- FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory ComputingChao Li, Zhicheng Xu, Bo Wen, Ruibin Mao et al.DAC 2024 · 6 citations
- Compact and Efficient CAM Architecture through Combinatorial Encoding and Self-Terminating Searching for In-Memory-Searching AcceleratorWeikai Xu, Jin Luo, Qianqian Huang, Ru HuangDAC 2024 · 1 citation
- Energy Efficient Dual Designs of FeFET-Based Analog In-Memory Computing with Inherent Shift-Add CapabilityZeyu Yang, Qingrong Huang, Yu Qian, Kai Ni et al.DAC 2024 · 9 citations
