Reshape and Adapt for Output Quantization (RAOQ): Quantization-aware Training for In-memory Computing Systems
Bonan Zhang, Chia-Yu Chen, Naveen Verma
Abstract
In-memory computing (IMC) has emerged as a promising solution to address both computation and data-movement challenges, by performing computation on data in-place directly in the memory array. IMC typically relies on analog operation, which makes analog-to-digital converters (ADCs) necessary, for converting results back to the digital domain. However, ADCs maintain computational efficiency by having limited precision, leading to substantial quantization errors in compute outputs. This work proposes RAOQ (Reshape and Adapt for Output Quantization) to overcome this issue, which comprises two classes of mechanisms including: 1) mitigating ADC quantization error by adjusting the statistics of activations and weights, through an activationshifting approach (A-shift) and a weight reshaping technique (W-reshape); 2) adapting AI models to better tolerate ADC quantization through a bit augmentation method (BitAug), complemented by the introduction of ADC-LoRA, a low-rank approximation technique, to reduce the training overhead. RAOQ demonstrates consistently high performance across different scales and domains of neural network models for computer vision and natural language processing (NLP) tasks at various bit precisions, achieving state-of-the-art results with practical IMC implementations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 967da8b6-2800-4cca-8ed6-114fccb2ee95Cited by top-tier papers2
- Analog In-memory Training on General Non-ideal Resistive Elements: The Impact of Response FunctionsZhaoxian Wu, Quan Xiao, Tayfun Gokmen, Omobayode Fagbohungbe et al.NeurIPS 2025 · 8 citations
- Analog Foundation ModelsJulian Büchel, Iason Chalas, Giovanni Acampa, An Chen et al.NeurIPS 2025 · 7 citations
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language ModelsYixiao Li, Yifan Yu, Chen Liang, Nikos Karampatziakis et al.ICLR 2024 · 217 citations
- QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language ModelsYuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen et al.ICLR 2024 · 179 citations
Related papers
- RILQ: Rank-Insensitive LoRA-Based Quantization Error Compensation for Boosting 2-Bit Large Language Model AccuracyGeonho Lee, Janghwan Lee, Sukjin Hong, Minsoo Kim et al.AAAI 2025 · 7 citations
- Leveraging Noise and Aggressive Quantization of In-Memory Computing for Robust DNN Hardware Against Adversarial Input and Weight AttacksSai Kiran Cherupally, Adnan Siraj Rakin, Shihui Yin, Mingoo Seok et al.DAC 2021 · 10 citations
- A Charge-Sharing based 8T SRAM In-Memory Computing for Edge DNN AccelerationKyeongho Lee, Sungsoo Cheon, Joongho Jo, Woong Choi et al.DAC 2021 · 34 citations
- YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AIZihao Xuan, Yuxuan Yang, Wei Xuan, Zijia Su et al.DAC 2025 · 1 citation
- SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM QuantizationYeonsik Park, Hyeonseong Kim, Seungkyu ChoiICLR 2026 · 1 citation
