InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching
Yilun Wang, Pengfei Chen, Haiyu Huang, Zilong He, Gou Tan, Chuanfu Zhang, Jingkai He, Zibin Zheng
Abstract
Modern software systems generate massive volumes of runtime logs, necessitating efficient and accurate log parsing to enable critical downstream tasks such as anomaly detection and root cause analysis. Recently, large language models (LLMs) have achieved advanced accuracy on log parsing, but their deployment in production environments faces two major limitations: First, the privacy risks associated with commercial LLMs, driving the adoption of local deployment. Second, online log parsing poses stringent challenges for latency and throughput. While recent methods reduce the number of LLM queries, they overlook the inherent overhead of LLM inference where concurrent log parsing requests can lead to performance degradation in the LLM inference system.
In this study, we present InferLog, the first LLM inference optimization method for online log parsing. Our key insight is that the inference efficiency emerges as the vital bottleneck in LLMbased online log parsing, rather than parsing accuracy. InferLog accelerates inference by designing (i) A prefix-aware ICL refinement strategy to refine the examples and permutation of in-context learning to improve the prefix caching efficiency. (ii) A rapid and task-specific configuration tuning pipeline based on meta-learning to find the optimal LLM inference system configuration. The experimental results based on Loghub-2k dataset and vLLM demonstrate that InferLog significantly outperforms existing inference optimization methods and markedly accelerates the state-of-the-art LLM-based log parsers without compromising parsing accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
Related papers
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language ModelsZeyang Ma, Dong Jae Kim, Tse-Hsun Peter ChenICSE 2025 · 7 citations
- MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context LearningJianbo Yu, Yixuan Li, Hai Xu, Kang Xu et al.AAAI 2026
- EPAS: Efficient Online Log Parsing via Asynchronous Scheduling of LLM QueriesXiaolei Chen, Jia Chen, Jie Shi, Peng Wang et al.ICDE 2025 · 3 citations
- LogParser-LLM: Advancing Efficient Log Parsing with Large Language ModelsAoxiao Zhong, Dengyao Mo, Guiyang Liu, Jinbu Liu et al.KDD 2024 · 41 citations
- VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingJinrui Sun, Tong Jia, Minghua He, Ying LiWWW 2026
