LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems
Wei Liu, Yongchao He, Bohan Zhao, Hongyi Wang, Zhenhua Li, Junping Zhao
2026Year
Abstract
Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in this context, as they are either too coarse-grained to capture performance bottlenecks, or too intrusive to avoid significant performance degradation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6ef0109d-d796-48df-bada-b4226b1941b4Related papers
- A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with CrystalLLMShaoke Xi, ChonLam Lao, Boyi Jia, Jiaqi Gao et al.SOSP 2026
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani et al.NeurIPS 2023 · 14 citations
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 17 citations
- MELTing Point: Mobile Evaluation of Language TransformersStefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed HaddadiMobiCom 2024 · 32 citations
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 4 citations
