Robust Methods for Developer Screening in Rapidly Evolving AI Contexts
Raphael Serafini, Nino Weber, Asli Yardim, Stefan Albert Horstmann, Alena Naiakshina
Abstract
The rise of AI-powered tools like ChatGPT enables non-programmers to bypass programming screening questions, undermining internal validity in usable security and privacy, and software engineering studies. Past ChatGPT-resistant tasks proposed static visual questions, which ChatGPT can now circumvent. Therefore, we tested alternative approaches such as video- and audio-based screeners that reveal key information step by step under strict time constraints to distinguish programmers from non-programmers. To this end, we conducted a study with 74 participants across three groups: programmers, non-programmers without AI assistance, and non-programmers using ChatGPT. Our results showed that audio-based screeners were robust against ChatGPT-based cheating, as non-programmers struggled to find correct answers within time limits, whereas programmers demonstrated high accuracy with minimal time pressure. Based on our findings, we recommend six audio-based ChatGPT-resistant screening questions that maximize screening effectiveness and efficiency and suggest a 215-second instrument that includes 95.87% of programmers while excluding 99.69% of non-programmers.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 19ed88cb-b31e-4bfa-9b18-183fe839a08eRelated papers
- ChatGPT-Resistant Screening Instrument for Identifying Non-ProgrammersRaphael Serafini, Clemens Otto, Stefan Albert Horstmann, Alena NaiakshinaICSE 2024 · 4 citations
- Do you really code? Designing and Evaluating Screening Questions for Online Surveys with ProgrammersAnastasia Danilova, Alena Naiakshina, Stefan Horstmann, Matthew SmithICSE 2021 · 5 citations
- Weak Programmers Need Not Apply, LLMs Welcome! Survey Screening in the AI EraIta Ryan, Utz Roedig, Klaas-Jan StolICSE 2026
- Impeding LLM-assisted Cheating in Introductory Programming Assignments via Adversarial PerturbationSaiful Islam Salim, Rubin Yuchan Yang, Alexander Cooper, Suryashree Ray et al.EMNLP 2024 · 4 citations
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 3 citations
