InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching
Yilun Wang, Pengfei Chen, Haiyu Huang, Zilong He, Gou Tan, Chuanfu Zhang, Jingkai He, Zibin Zheng
摘要
Modern software systems generate massive volumes of runtime logs, necessitating efficient and accurate log parsing to enable critical downstream tasks such as anomaly detection and root cause analysis. Recently, large language models (LLMs) have achieved advanced accuracy on log parsing, but their deployment in production environments faces two major limitations: First, the privacy risks associated with commercial LLMs, driving the adoption of local deployment. Second, online log parsing poses stringent challenges for latency and throughput. While recent methods reduce the number of LLM queries, they overlook the inherent overhead of LLM inference where concurrent log parsing requests can lead to performance degradation in the LLM inference system.
In this study, we present InferLog, the first LLM inference optimization method for online log parsing. Our key insight is that the inference efficiency emerges as the vital bottleneck in LLMbased online log parsing, rather than parsing accuracy. InferLog accelerates inference by designing (i) A prefix-aware ICL refinement strategy to refine the examples and permutation of in-context learning to improve the prefix caching efficiency. (ii) A rapid and task-specific configuration tuning pipeline based on meta-learning to find the optimal LLM inference system configuration. The experimental results based on Loghub-2k dataset and vLLM demonstrate that InferLog significantly outperforms existing inference optimization methods and markedly accelerates the state-of-the-art LLM-based log parsers without compromising parsing accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
相关 Paper
- LibreLog: Accurate and Efficient Unsupervised Log Parsing Using Open-Source Large Language ModelsZeyang Ma, Dong Jae Kim, Tse-Hsun Peter ChenICSE 2025 · 被引用 7 次
- MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context LearningJianbo Yu, Yixuan Li, Hai Xu, Kang Xu 等AAAI 2026
- EPAS: Efficient Online Log Parsing via Asynchronous Scheduling of LLM QueriesXiaolei Chen, Jia Chen, Jie Shi, Peng Wang 等ICDE 2025 · 被引用 3 次
- LogParser-LLM: Advancing Efficient Log Parsing with Large Language ModelsAoxiao Zhong, Dengyao Mo, Guiyang Liu, Jinbu Liu 等KDD 2024 · 被引用 41 次
- VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingJinrui Sun, Tong Jia, Minghua He, Ying LiWWW 2026
