ChatGPT-Resistant Screening Instrument for Identifying Non-Programmers
Raphael Serafini, Clemens Otto, Stefan Albert Horstmann, Alena Naiakshina
摘要
To ensure the validity of software engineering and IT security studies with professional programmers, it is essential to identify participants without programming skills. Existing screening questions are efficient, cheating robust, and effectively differentiate programmers from non-programmers. However, the release of ChatGPT raises concerns about their continued effectiveness in identifying non-programmers. In a simulated attack, we showed that Chat-GPT can easily solve existing screening questions. Therefore, we designed new ChatGPT-resistant screening questions using visual concepts and code comprehension tasks. We evaluated 28 screening questions in an online study with 121 participants involving programmers and non-programmers. Our results showed that questions using visualizations of well-known programming concepts performed best in differentiating between programmers and nonprogrammers. Participants prompted to use ChatGPT struggled to solve the tasks. They considered ChatGPT ineffective and changed their strategy after a few screening questions. In total, we present six ChatGPT-resistant screening questions that effectively identify non-programmers. We provide recommendations on setting up a ChatGPT-resistant screening instrument that takes less than three minutes to complete by excluding 99.47% of non-programmers while including 94.83% of programmers.
• Human-centered computing → Empirical studies in HCI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- You Get Where You're Looking for: The Impact of Information Sources on Code SecurityYasemin Acar, Michael Backes, Sascha Fahl, Doowon Kim 等S&P 2016 · 被引用 325 次
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 被引用 133 次
- "I Have No Idea What I'm Doing" - On the Usability of Deploying HTTPSKatharina Krombholz, Wilfried Mayer, Martin Schmiedecker, Edgar R. WeipplUSENIX Security 2017 · 被引用 114 次
- Recruiting Participants With Programming Skills: A Comparison of Four Crowdsourcing Platforms and a CS Student Mailing ListMohammad Tahaei, Kami VanieaCHI 2022 · 被引用 45 次
相关 Paper
- Robust Methods for Developer Screening in Rapidly Evolving AI ContextsRaphael Serafini, Nino Weber, Asli Yardim, Stefan Albert Horstmann 等CHI 2026 · 被引用 1 次
- Testing Time Limits in Screener Questions for Online Surveys with ProgrammersAnastasia Danilova, Stefan Horstmann, Matthew Smith, Alena NaiakshinaICSE 2022 · 被引用 8 次
- Do you really code? Designing and Evaluating Screening Questions for Online Surveys with ProgrammersAnastasia Danilova, Alena Naiakshina, Stefan Horstmann, Matthew SmithICSE 2021 · 被引用 5 次
- Weak Programmers Need Not Apply, LLMs Welcome! Survey Screening in the AI EraIta Ryan, Utz Roedig, Klaas-Jan StolICSE 2026
- Using AI Assistants in Software Development: A Qualitative Study on Security Practices and ConcernsJan H. Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden 等CCS 2024 · 被引用 14 次
