AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUs
Shiheng Cao, Junmin Wu, Junshi Chen, Hong An, Zhibin Yu
摘要
Large language model (LLM) inference applications are surging in recent years, which largely relies on modern GPUs.On the other hand, GPU analytical model is a commonly used tool for architects to precisely identify bottlenecks quickly with deep insights.However, existing GPU analytical models fall short of accurately modeling LLM inference applications on modern GPUs, because of unsuitable tensor core modeling, ignoring constant cache as well as instruction cache modeling and abstracting away important details for LLM inference applications.To address this problem, we propose a novel analytical model dubbed AMALI to accurately model LLM inference on modern GPUs with three innovations.First, we develop an instruction modifier and throughput based tensor core model by accurately capturing the math pipe throttle stalls to enhance the architecture modeling for modern GPUs.Second, we propose analytical models for constant cache and instruction cache by developing micro-benchmarks to measure CUDA kernel launching latencies.This significantly improves AMALI's accuracy compared to real GPU hardware.Finally, we design a multi-warp model by leveraging warp instruction number distribution to reflect LLM inference application characteristics.We validate AMALI on an A100 GPU by using typical LLM inference applications.The results show that AMALI reduces the MAPE (mean absolute percentage error) from 127.56% to 23.59%
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo 等ISCA 2022 · 被引用 25 次
- MDM: The GPU Memory Divergence ModelLu Wang, Magnus Jahre, Almutaz Adileh, Lieven EeckhoutMICRO 2020 · 被引用 27 次
- LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationHyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang 等HPCA 2026 · 被引用 1 次
- MPK: A Compiler and Runtime for Mega-Kernelizing Tensor ProgramsXinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji 等OSDI 2026 · 被引用 20 次
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma 等ICML 2026 · 被引用 7 次
