What Makes Code Generation Ethically Sourced?
Zhuolin Xu, Chenglin Li, Qiushi Li, Shin Hwei Tan
Abstract
Several code generation models have been proposed to help reduce time and effort in solving software-related tasks. To ensure responsible AI, there are growing interests over various ethical issues (e.g., unclear licensing, privacy, fairness, and environment impact). These studies have the overarching goal of ensuring ethically sourced generation, which has gained growing attention in speech synthesis and image generation. In this paper, we introduce the novel notion of Ethically Sourced Code Generation (ES-CodeGen) to refer to managing all processes involved in code generation model development from data collection to post-deployment via ethical and sustainable practices. To build a taxonomy of ES-CodeGen, we perform a two-phase literature review where we reviewed 803 papers across various domains and specific to AI-based code generation. We identified 71 relevant papers with 10 initial dimensions of ES-CodeGen. To refine our dimensions and gain insights on consequences of ES-CodeGen, we surveyed 32 practitioners, which include six developers who submitted GitHub issues to opt-out from the Stack dataset (these impacted users have real-world experience of ethically sourced issues in code generation models). The results lead to 11 dimensions of ES-CodeGen with a new dimension on code quality as practitioners have noted its importance. We also identified consequences, artifacts, and stages relevant to ES-CodeGen. Our post-survey reflection showed that most practitioners tended to ignore society-related dimensions despite their importance. Most practitioners either agreed or strongly agreed that our survey help improve their understanding of ES-CodeGen. Our study calls for attention of various ethical issues towards ES-CodeGen.
• Software and its engineering → Software maintenance tools; • Human-centered computing → Empirical studies in collaborative and social computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- An Empirical Study on Software Bill of Materials: Where We Stand and the Road AheadBoming Xia, Tingting Bi, Zhenchang Xing, Qinghua Lu et al.ICSE 2023 · 82 citations
- Uncovering and Quantifying Social Biases in Code GenerationYan Liu, Xiaokang Chen, Yan Gao, Zhe Su et al.NeurIPS 2023 · 47 citations
- Using AI Assistants in Software Development: A Qualitative Study on Security Practices and ConcernsJan H. Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden et al.CCS 2024 · 14 citations
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu et al.ASE 2025 · 2 citations
- CodexLeaks: Privacy Leaks from Code Generation Language Models in GitHub CopilotLiang Niu, Muhammad Shujaat Mirza, Zayd Maradni, Christina PöpperUSENIX Security 2023
Related papers
- Do Users Write More Insecure Code with AI Assistants?Neil Perry, Megha Srivastava, Deepak Kumar, Dan BonehCCS 2023 · 150 citations
- A User-centered Security Evaluation of CopilotOwura Asare, Meiyappan Nagappan, N. AsokanICSE 2024 · 11 citations
- An Empirical Study of Code Clones from Commercial AI Code GeneratorsWeibin Wu, Haoxuan Hu, Zhaoji Fan, Yitong Qiao et al.FSE 2025 · 2 citations
- CodeIPPrompt: Intellectual Property Infringement Assessment of Code Language ModelsZhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang et al.ICML 2023 · 42 citations
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 6 citations
