Weak Programmers Need Not Apply, LLMs Welcome! Survey Screening in the AI Era
Ita Ryan, Utz Roedig, Klaas-Jan Stol
Abstract
Quantitative research on software development practice often depends on the results of anonymous online surveys. However, it is difficult to verify that the respondents to such surveys are genuine developers, especially if survey responses are paid for, incentivising respondents to game the system and attracting the attention of bots and professional survey farms. One solution is to include screening questions to assess respondents’ coding skills, but this introduces the risk of unfairly excluding considered responses from developers who spent time and energy on the survey. We reviewed the responses of 86 developers who were excluded via screening from analysis of a large unpaid developer survey (n=1,048). We estimated that up to 86% of the developers who were excluded were genuine developers. The advent of large language models (LLMs) casts further doubt on the use of screening questions. Investigating LLM ability, we found that five sample LLMs could answer most widely-used screening questions. Powerful LLM-based survey-taking tools now exist. We researched current survey screening techniques and found that they are susceptible to fraud via LLM use, bots and survey farms. Recruitment strategy may be a better screening technique. We recommend that if survey respondents are compensated only known developers should be invited. Surveys that are widely distributed, e.g. on social media, should not compensate respondents. Instead, researchers should focus on good survey design and sustained, imaginative recruitment strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a16a6678-beaa-4ecc-b873-74064a3912eeBuilds on14
- A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and ChallengesJenny T. Liang, Chenyang Yang, Brad A. MyersICSE 2024 · 126 citations
- Recruiting Participants With Programming Skills: A Comparison of Four Crowdsourcing Platforms and a CS Student Mailing ListMohammad Tahaei, Kami VanieaCHI 2022 · 45 citations
- Building and Validating a Scale for Secure Software Development Self-EfficacyDaniel Votipka, Desiree Abrokwa, Michelle L. MazurekCHI 2020 · 35 citations
- One size does not fit all: a grounded theory and online survey study of developer preferences for security warning typesAnastasia Danilova, Alena Naiakshina, Matthew SmithICSE 2020 · 24 citations
- Measuring Secure Coding Practice and Culture: A Finger Pointing at the Moon is not the MoonIta Ryan, Utz Roedig, Klaas-Jan StolICSE 2023 · 13 citations
Related papers
- Do you really code? Designing and Evaluating Screening Questions for Online Surveys with ProgrammersAnastasia Danilova, Alena Naiakshina, Stefan Horstmann, Matthew SmithICSE 2021 · 5 citations
- Simulacrum of Stories: Examining Large Language Models as Qualitative Research ParticipantsShivani Kapania, William Agnew, Motahhare Eslami, Hoda Heidari et al.CHI 2025 · 59 citations
- Safeguarding Crowdsourcing Surveys from ChatGPT through Prompt InjectionChaofan Wang, Samuel Kernan Freire, Mo Zhang, Jing Wei et al.CSCW 2025 · 1 citation
- Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case StudyPerttu Hämäläinen, Mikke Tavast, Anton KunnariCHI 2023 · 244 citations
- Deception at Scale: Deceptive Designs in 1K LLM-Generated E-Commerce ComponentsZiwei Chen, Jiawen Shen, Luna, Hanyu Zhang et al.CHI 2026 · 3 citations
