Do Hackers Dream of Electric Teachers?: A Large-Scale, In-Situ Measurement of Cybersecurity Student Behaviors and Educational Performance with AI Tutors
Michael Tompkins, Nihaarika Agarwal, Ananta Soneji, Robert Wasinger, Connor Nelson, Kevin Leach, Rakibul Hasan, Adam Doupé, Daniel Votipka, Yan Shoshitaishvili, Jaron Mink
Abstract
To meet the ever-increasing demands of the cybersecurity workforce, AI tutors have been proposed for personalized, scalable education. But, while AI tutors have shown promise in introductory programming courses, no work has evaluated their use in hands-on exploration and exploitation exercises (e.g.,"Capture the Flag") commonly used to teach cybersecurity. In particular, it is unclear how students use AI tutors, or what types of use correlate with greater success in solving the challenges in real, large-scale cybersecurity courses. To answer this, we conducted a semester-long observational study of an embedded AI tutor with 309 students in an upper-division introductory cybersecurity course. By analyzing 142,526 student queries sent to the AI tutor across 383 cybersecurity challenges spanning 9 core cybersecurity topics and an accompanying end-of-semester survey, we find (1) what queries and conversation styles students use with AI tutors, (2) how these styles relate to challenge completion, and (3) students'perceptions of AI tutors in cybersecurity education. In particular, we identify three broad AI tutor conversation styles among students: Short (bounded, few-turn exchanges), Reactive (repeatedly submitting code and errors), and Proactive (driving problem-solving through targeted inquiry). We also find that these styles are significantly correlated with challenge completion, and that the completion-rate gap between styles widens as materials become more advanced. Furthermore, students valued the tutor's availability but reported that it became less useful for harder material. Based on our results, we provide suggestions for security educators and developers on practical AI tutor use.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator NeedsMajeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley et al.CHI 2024 · 246 citations
- Hackers vs. Testers: A Comparison of Software Vulnerability Discovery ProcessesDaniel Votipka, Rock Stevens, Elissa M. Redmiles, Jeremy Hu et al.S&P 2018 · 151 citations
- HackEd: A Pedagogical Analysis of Online Vulnerability Discovery ExercisesDaniel Votipka, Eric Zhang, Michelle L. MazurekS&P 2021 · 22 citations
- Vulnerability Discovery for All: Experiences of Marginalization in Vulnerability DiscoveryKelsey R. Fulton, Samantha Katcher, Kevin Song, Marshini Chetty et al.S&P 2023
- "I'm trying to learn...and I'm shooting myself in the foot": Beginners' Struggles When Solving Binary Exploitation ExercisesJames Mattei, Christopher Pellegrini, Matthew Soto, Marina Sanusi Bohuk et al.USENIX Security 2025
Related papers
- From Code Generation to Conceptual Learning: Student Use of LLMs in a Web Programming CourseHajara-Yasmin Isa, Matthew Weston, Muhammad Rizky Wellyanto, Ishita Karna et al.CHI 2026 · 1 citation
- From Assistance to Autonomy: An Empirical Study of AI Use in a Live Capture-the-Flag (CTF) CompetitionTingxuan Tang, Nicolas Janis, Kalyn Asher Montague, Kevin Eykholt et al.USENIX Security 2026
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration TestingJustin W. Lin, Eliot Jones, Donovan Jasper, Ethan Ho et al.ICLR 2026 · 18 citations
- How Do Programming Students Use Generative AI?Christian Rahe, Walid MaalejFSE 2025 · 12 citations
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
