Robust Methods for Developer Screening in Rapidly Evolving AI Contexts
Raphael Serafini, Nino Weber, Asli Yardim, Stefan Albert Horstmann, Alena Naiakshina
摘要
The rise of AI-powered tools like ChatGPT enables non-programmers to bypass programming screening questions, undermining internal validity in usable security and privacy, and software engineering studies. Past ChatGPT-resistant tasks proposed static visual questions, which ChatGPT can now circumvent. Therefore, we tested alternative approaches such as video- and audio-based screeners that reveal key information step by step under strict time constraints to distinguish programmers from non-programmers. To this end, we conducted a study with 74 participants across three groups: programmers, non-programmers without AI assistance, and non-programmers using ChatGPT. Our results showed that audio-based screeners were robust against ChatGPT-based cheating, as non-programmers struggled to find correct answers within time limits, whereas programmers demonstrated high accuracy with minimal time pressure. Based on our findings, we recommend six audio-based ChatGPT-resistant screening questions that maximize screening effectiveness and efficiency and suggest a 215-second instrument that includes 95.87% of programmers while excluding 99.69% of non-programmers.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ChatGPT-Resistant Screening Instrument for Identifying Non-ProgrammersRaphael Serafini, Clemens Otto, Stefan Albert Horstmann, Alena NaiakshinaICSE 2024 · 被引用 4 次
- Do you really code? Designing and Evaluating Screening Questions for Online Surveys with ProgrammersAnastasia Danilova, Alena Naiakshina, Stefan Horstmann, Matthew SmithICSE 2021 · 被引用 5 次
- Weak Programmers Need Not Apply, LLMs Welcome! Survey Screening in the AI EraIta Ryan, Utz Roedig, Klaas-Jan StolICSE 2026
- Impeding LLM-assisted Cheating in Introductory Programming Assignments via Adversarial PerturbationSaiful Islam Salim, Rubin Yuchan Yang, Alexander Cooper, Suryashree Ray 等EMNLP 2024 · 被引用 4 次
- AI Wrote My Paper and All I Got was This False Negative:* Measuring the Efficacy of Commercial AI Text DetectorsSeth Layton, Bernardo B. P. Medeiros, Kevin R. B. Butler, Patrick TraynorS&P 2026 · 被引用 3 次
