Lune

ICSE2026顶会

InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix Caching

Yilun Wang, Pengfei Chen, Haiyu Huang, Zilong He, Gou Tan, Chuanfu Zhang, Jingkai He, Zibin Zheng

2026年份
1被引次数

摘要

Modern software systems generate massive volumes of runtime logs, necessitating efficient and accurate log parsing to enable critical downstream tasks such as anomaly detection and root cause analysis. Recently, large language models (LLMs) have achieved advanced accuracy on log parsing, but their deployment in production environments faces two major limitations: First, the privacy risks associated with commercial LLMs, driving the adoption of local deployment. Second, online log parsing poses stringent challenges for latency and throughput. While recent methods reduce the number of LLM queries, they overlook the inherent overhead of LLM inference where concurrent log parsing requests can lead to performance degradation in the LLM inference system.

In this study, we present InferLog, the first LLM inference optimization method for online log parsing. Our key insight is that the inference efficiency emerges as the vital bottleneck in LLMbased online log parsing, rather than parsing accuracy. InferLog accelerates inference by designing (i) A prefix-aware ICL refinement strategy to refine the examples and permutation of in-context learning to improve the prefix caching efficiency. (ii) A rapid and task-specific configuration tuning pipeline based on meta-learning to find the optimal LLM inference system configuration. The experimental results based on Loghub-2k dataset and vLLM demonstrate that InferLog significantly outperforms existing inference optimization methods and markedly accelerates the state-of-the-art LLM-based log parsers without compromising parsing accuracy.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖