AMALI: An Analytical Model for Accurately Modeling LLM Inference on Modern GPUs
Shiheng Cao, Junmin Wu, Junshi Chen, Hong An, Zhibin Yu
Abstract
Large language model (LLM) inference applications are surging in recent years, which largely relies on modern GPUs.On the other hand, GPU analytical model is a commonly used tool for architects to precisely identify bottlenecks quickly with deep insights.However, existing GPU analytical models fall short of accurately modeling LLM inference applications on modern GPUs, because of unsuitable tensor core modeling, ignoring constant cache as well as instruction cache modeling and abstracting away important details for LLM inference applications.To address this problem, we propose a novel analytical model dubbed AMALI to accurately model LLM inference on modern GPUs with three innovations.First, we develop an instruction modifier and throughput based tensor core model by accurately capturing the math pipe throttle stalls to enhance the architecture modeling for modern GPUs.Second, we propose analytical models for constant cache and instruction cache by developing micro-benchmarks to measure CUDA kernel launching latencies.This significantly improves AMALI's accuracy compared to real GPU hardware.Finally, we design a multi-warp model by leveraging warp instruction number distribution to reflect LLM inference application characteristics.We validate AMALI on an A100 GPU by using typical LLM inference applications.The results show that AMALI reduces the MAPE (mean absolute percentage error) from 127.56% to 23.59%
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aec3cb0e-eab7-4905-a292-89e3af86c48eCited by top-tier papers1
Ask how each one uses itRelated papers
- GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsJounghoo Lee, Yeonan Ha, Suhyun Lee, Jinyoung Woo et al.ISCA 2022 · 25 citations
- MDM: The GPU Memory Divergence ModelLu Wang, Magnus Jahre, Almutaz Adileh, Lieven EeckhoutMICRO 2020 · 27 citations
- LiLo: Harnessing the on-Chip Accelerators in Intel CPUs for Compressed LLM Inference AccelerationHyungyo Kim, Qirong Xia, Jinghan Huang, Nachuan Wang et al.HPCA 2026 · 1 citation
- MPK: A Compiler and Runtime for Mega-Kernelizing Tensor ProgramsXinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji et al.OSDI 2026 · 20 citations
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma et al.ICML 2026 · 7 citations
