Xpert: Empowering Incident Management with Query Recommendations via Large Language Models
Yuxuan Jiang, Chaoyun Zhang, Shilin He, Zhihao Yang, Minghua Ma, Si Qin, Yu Kang, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang
Abstract
Large-scale cloud systems play a pivotal role in modern IT infrastructure. However, incidents occurring within these systems can lead to service disruptions and adversely affect user experience. To swiftly resolve such incidents, on-call engineers depend on crafting domain-specific language (DSL) queries to analyze telemetry data. However, writing these queries can be challenging and timeconsuming. This paper presents a thorough empirical study on the utilization of queries of KQL, a DSL employed for incident management in a large-scale cloud management system at Microsoft. The findings obtained underscore the importance and viability of KQL queries recommendation to enhance incident management. Building upon these valuable insights, we introduce Xpert, an end-to-end machine learning framework that automates KQL recommendation process. By leveraging historical incident data and large language models, Xpert generates customized KQL queries tailored to new incidents. Furthermore, Xpert incorporates a novel performance metric called Xcore, enabling a thorough evaluation of query quality from three comprehensive perspectives. We conduct extensive evaluations of Xpert, demonstrating its effectiveness in offline settings. Notably, we deploy Xpert in the real production environment of a large-scale incident management system in Microsoft, validating its efficiency in supporting incident management. To the best of our knowledge, this paper represents the first empirical study of its kind, and Xpert stands as a pioneering DSL query recommendation framework designed for incident management.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 109ff518-e8ac-4c11-8980-eb3e26162beeCited by top-tier papers10
- Revisiting VAE for Unsupervised Time Series Anomaly Detection: A Frequency PerspectiveZexin Wang, Changhua Pei, Minghua Ma, Xin Wang et al.WWW 2024 · 90 citations
- STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern CloudsYinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su et al.NeurIPS 2025 · 35 citations
- CUPID: Improving Battle Fairness and Position Satisfaction in Online MOBA Games with a Re-matchmaking SystemGe Fan, Chaoyun Zhang, Kai Wang, Yingjie Li et al.CSCW 2024 · 7 citations
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu et al.ICML 2026 · 4 citations
- FAIL: Analyzing Software Failures from the News Using LLMsDharun Anandayuvaraj, Matthew Campbell, Arav Tewari, James C. DavisASE 2024 · 3 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
Related papers
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- StepFly: Agentic Troubleshooting Guide Automation for Incident DiagnosisJiayi Mao, Liqun Li, Yanjie Gao, Zegang Peng et al.FSE 2026 · 1 citation
- LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call AutomationRuowei Fu, Yang Zhang, Zeyu Che, Xin Wu et al.ASE 2025
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Towards LLM-Based Failure Localization in Production-Scale NetworksChenxu Wang, Xumiao Zhang, Runwei Lu, Xianshang Lin et al.SIGCOMM 2025 · 13 citations
