Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
Xilong Xie, Liang Wang, Limin Xiao, Meng Han, Lei Liu, Xiangrong Xu, Jinquan Wang, Zhen Song, Xiaojian Liao
Abstract
The quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ca0d9770-bbf5-4499-b1c5-10547fbafd9bCited by top-tier papers3
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu et al.ISCA 2026 · 4 citations
- ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM InferenceXiongwei Zhu, Xiaojian Liao, Tianyang Jiang, Yusen Zhang et al.ICML 2026 · 2 citations
- Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative DecodingSoongyu Choi, Yuntae Kim, Muyoung Son, Joo-Young KimISCA 2026
Related papers
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang et al.HPCA 2025 · 23 citations
- Oltron: Algorithm-Hardware Co-design for Outlier-Aware Quantization of LLMs with Inter-/Intra-Layer AdaptationChenhao Xue, Chen Zhang, Xun Jiang, Zhutianya Gao et al.DAC 2024 · 11 citations
- OutlierCIM: Outlier-Aware Digital CIM-Based LLM Accelerator with Hybrid-Strategy Quantization and Unified FP-INT ComputationZihan Zou, Shikuang Chen, Chen Zhang, Xing Wang et al.DAC 2025
- OPAL: Outlier-Preserved Microscaling Quantization Accelerator for Generative Large Language ModelsJahyun Koo, Dahoon Park, Sangwoo Jung, Jaeha KungDAC 2024 · 12 citations
- DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs AccelerationZhuoquan Yu, Huidong Ji, Yue Cao, Junfu Wu et al.DAC 2025 · 1 citation
