Lune

CCS2026Top-tier venue

PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation

Yixuan Liu, Zilong Zhen, Yin Wu, Yi Li

2026Year

Abstract

As Large Language Model (LLM) agents increasingly automate offensive operations across the cyber kill chain, their efficacy in complex local post-exploitation tasks remains inadequately quantified. Among these, Linux privilege escalation is a key step between initial access and full system compromise. However, existing evaluations for this task are limited by small sample sizes (<15 scenarios), lacking the scale to compare model capabilities under executable verification. To address this, we present PrivEscalate, a large-scale benchmark for Linux privilege escalation, comprising 531 Dockerized scenarios spanning 14 distinct sub-categories. To measure sensitivity to environmental distractors, we additionally derive 329 parameterized variant scenarios so that each model's demonstrated successes can be retested under matched perturbations.

Evaluating six LLMs across three agent architectures reveals: (i) Model capability is heterogeneous across vulnerability classes, with no single model dominating across the high-prevalence classes, motivating multi-dimensional risk assessments. (ii) LLM successes are sensitive to environmental perturbation, so configuration rotation can disrupt some exploit attempts but does not eliminate the measured risk. (iii) Agent architectures can materially change success rates and reorder model rankings, though the magnitude is model-dependent. Leveraging these insights, we develop PrivEscAgent, a domain-specialized wrapper that augments a generic Re-Act agent with deterministic enumeration, category matching, and step planning. PrivEscAgent improves over prior Linux privilegeescalation agent baselines without underlying LLM modifications. We release PrivEscalate as an open-source, Dockerized measurement instrument supporting both LLM agent evaluation and broader Linux privilege escalation research, including defensive tool validation and red-team training.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bd0fafe6-1a87-4c4b-8cb5-1f6e914f6894

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines