ADAngel: Accelerating Arbitrary-Precision Quantized LLMs with Adaptive Computing Mapping
Yao Liu, Wenjie Wang, Yifei Feng, Bo Peng, Jianguo Yao, Haibing Guan
摘要
Arbitrary-Precision Quantization (APQ), which uses asymmetric bit-widths for weights and activations (e.g., W4A8), is a prevalent technique for LLM inference because of its excellent accuracy-performance balance. APQ transforms the general matrix multiplications (GEMM), the core of LLM computation, into mixed-precision GEMM (mpGEMM) whose two operand matrices have different quantization bit-widths. However, we identify that the computation paradigms of mpGEMM in current APQ LLM inference systems are sub-optimal because the shapes and bit-widths of mpGEMM tasks in APQ LLM are highly variable, whereas existing static and workload-unaware paradigms can only accelerate mpGEMM tasks with the same or similar shapes and bit-widths. Based on this finding, we propose ADAngel, a framework for creating a workload-adaptive mpGEMM computation core for target LLMs. The theoretical foundation of ADAngel is the DPR (Decomposition-Partial Product-Reconstruction) computation model, which enables systematic generation of a diverse portfolio of mpGEMM algorithms by specifying different bit-partition schemes. Guided by this model, ADAngel constructs a Computation Strategy Set comprising several highly optimized mpGEMM kernels, and exhaustively analyzes the strategy set to create an Oracle Policy Map, which enables a lightweight dispatcher to select and execute the optimal kernel for runtime mpGEMM tasks with negligible overhead. Our evaluation shows that the ADAngel-specialized engine achieves up to a 5.10× speedup in decode throughput over llama.cpp; while in the prefill stage, it demonstrates its adaptivity by delivering speedups ranging from 1.17× to 2.38× over TensorRT-LLM in Time-To-First-Token (TTFT).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 被引用 503 次
- BiLLM: Pushing the Limit of Post-Training Quantization for LLMsWei Huang, Yangdong Liu, Haotong Qin, Ying Li 等ICML 2024 · 被引用 161 次
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee 等ICLR 2024 · 被引用 134 次
相关 Paper
- COMET: Towards Practical W4A4KV4 LLMs ServingLian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu 等ASPLOS 2025 · 被引用 5 次
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu 等AAAI 2025 · 被引用 24 次
- UniCore: A Bit-Width Scalable GEMM Unit for Unified LLM InferenceYonghao Chen, Jiaxiang Zou, Xingyu Chen, Chenxi Xu 等ISCA 2026
- LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM ServingHuanqi Hu, Bowen Xiao, Shixuan Sun, Jianian Yin 等SC 2025 · 被引用 2 次
- MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language ModelsWenyuan Liu, Haoqian Meng, Yilun Luo, Peng Zhang 等ICLR 2026 · 被引用 12 次
