PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation
Yixuan Liu, Zilong Zhen, Yin Wu, Yi Li
摘要
As Large Language Model (LLM) agents increasingly automate offensive operations across the cyber kill chain, their efficacy in complex local post-exploitation tasks remains inadequately quantified. Among these, Linux privilege escalation is a key step between initial access and full system compromise. However, existing evaluations for this task are limited by small sample sizes (<15 scenarios), lacking the scale to compare model capabilities under executable verification. To address this, we present PrivEscalate, a large-scale benchmark for Linux privilege escalation, comprising 531 Dockerized scenarios spanning 14 distinct sub-categories. To measure sensitivity to environmental distractors, we additionally derive 329 parameterized variant scenarios so that each model's demonstrated successes can be retested under matched perturbations.
Evaluating six LLMs across three agent architectures reveals: (i) Model capability is heterogeneous across vulnerability classes, with no single model dominating across the high-prevalence classes, motivating multi-dimensional risk assessments. (ii) LLM successes are sensitive to environmental perturbation, so configuration rotation can disrupt some exploit attempts but does not eliminate the measured risk. (iii) Agent architectures can materially change success rates and reorder model rankings, though the magnitude is model-dependent. Leveraging these insights, we develop PrivEscAgent, a domain-specialized wrapper that augments a generic Re-Act agent with deterministic enumeration, category matching, and step planning. PrivEscAgent improves over prior Linux privilegeescalation agent baselines without underlying LLM modifications. We release PrivEscalate as an open-source, Dockerized measurement instrument supporting both LLM agent evaluation and broader Linux privilege escalation research, including defensive tool validation and red-team training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration TestingGelei Deng, Yi Liu, Víctor Mayoral Vilches, Peng Liu 等USENIX Security 2024 · 被引用 186 次
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security TasksHwiwon Lee, Ziqi Zhang, Hanxiao Lu, Lingming ZhangNeurIPS 2025 · 被引用 86 次
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based AgentsYifu Guo, Jiaye Lin, Huacan Wang, Yuzhen Han 等NeurIPS 2025 · 被引用 73 次
- ChainReactor: Automated Privilege Escalation Chain Discovery via AI PlanningGiulio De Pasquale, Ilya Grishchenko, Riccardo Iesari, Gabriel Pizarro 等USENIX Security 2024 · 被引用 15 次
- SCAVY: Automated Discovery of Memory Corruption Targets in Linux Kernel for Privilege EscalationErin Avllazagaj, Yonghwi Kwon, Tudor DumitrasUSENIX Security 2024 · 被引用 7 次
相关 Paper
- PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation CapabilitiesZicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 等ICLR 2026 · 被引用 6 次
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application VulnerabilitiesYuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li 等ICML 2025 · 被引用 1 次
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis 等ICML 2026 · 被引用 9 次
- Detecting Privilege Escalation in Polyglot Microservices via Agentic Program AnalysisPenghui Li, Hong Yau Chong, Yinzhi Cao, Junfeng YangS&P 2026 · 被引用 3 次
- Training Language Model Agents to Find Vulnerabilities with CTF-DojoTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar 等ICML 2026 · 被引用 12 次
