From Assistance to Autonomy: An Empirical Study of AI Use in a Live Capture-the-Flag (CTF) Competition
Tingxuan Tang, Nicolas Janis, Kalyn Asher Montague, Kevin Eykholt, Dhilung Kirat, Youngja Park, Jiyong Jang, Adwait Nadkarni, Yue Xiao
摘要
Capture-the-Flag (CTF) competitions are increasingly becoming a testbed for evaluating AI capabilities at solving security tasks, due to their controlled environments and objective success criteria. Existing evaluations have focused on how successful AI is at solving individual CTF challenges in isolation from human CTF players. As AI usage increases in both academic and industrial settings, it is equally likely that human CTF players may collaborate with AI agents to solve CTF challenges. This possibility exposes a key knowledge gap: how do human players perceive AI CTF assistance; when assistance is provided, in what ways do they collaborate and is it effective with respect to human performance; how do humans assisted by AI compare to the performance of fully autonomous AI agents on the same set of challenges. We address this gap with the first empirical study of AI assistance in a live, onsite CTF. In a study with 41 participants (out of the total 95 that participated in the CTF), we qualitatively study (i) how participants' perception, trust, and expectations shift before versus after hands-on AI use, and (ii) how participants collaborate with an instrumented AI assistant. Moreover, we also (iii) benchmark four autonomous CTF agents on the same fresh challenge set to compare outcomes with human teams and analyze agent trajectories. We find that, for human players, AI literacy and domain knowledge are complementary competencies, and both have irreplaceable advantages. Proficient and efficient use of AI amplifies professional skills. Importantly, although advanced autonomous agents showed outstanding performance, human-in-the-loop is the winning paradigm where AI accelerates exploration while humans provide targeted guidance and verification. We conclude with implications for the future design of CTF competitions and for building effective human-in-the-loop AI systems for security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
- Exploring ChatGPT's Capabilities on Vulnerability ManagementPeiyu Liu, Junming Liu, Lirong Fu, Kangjie Lu 等USENIX Security 2024 · 被引用 52 次
- From One Thousand Pages of Specification to Unveiling Hidden Bugs: Large Language Model Assisted Fuzzing of Matter IoT DevicesXiaoyue Ma, Lannan Luo, Qiang ZengUSENIX Security 2024 · 被引用 49 次
- ChainReactor: Automated Privilege Escalation Chain Discovery via AI PlanningGiulio De Pasquale, Ilya Grishchenko, Riccardo Iesari, Gabriel Pizarro 等USENIX Security 2024 · 被引用 15 次
- Training Language Model Agents to Find Vulnerabilities with CTF-DojoTerry Yue Zhuo, Dingmin Wang, Hantian Ding, Varun Kumar 等ICML 2026 · 被引用 12 次
相关 Paper
- Measuring and Augmenting Large Language Models for Solving Capture-the-Flag ChallengesZimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma 等CCS 2025
- Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF BenchmarkMinghao Shao, Nanda Rani, Kimberly Milner, Haoran Xi 等AAAI 2026 · 被引用 5 次
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration TestingJustin W. Lin, Eliot Jones, Donovan Jasper, Ethan Ho 等ICLR 2026 · 被引用 18 次
- Do Hackers Dream of Electric Teachers?: A Large-Scale, In-Situ Measurement of Cybersecurity Student Behaviors and Educational Performance with AI TutorsMichael Tompkins, Nihaarika Agarwal, Ananta Soneji, Robert Wasinger 等CCS 2026
- EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security VulnerabilitiesTalor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret 等ICML 2025
