What Makes Code Generation Ethically Sourced?
Zhuolin Xu, Chenglin Li, Qiushi Li, Shin Hwei Tan
摘要
Several code generation models have been proposed to help reduce time and effort in solving software-related tasks. To ensure responsible AI, there are growing interests over various ethical issues (e.g., unclear licensing, privacy, fairness, and environment impact). These studies have the overarching goal of ensuring ethically sourced generation, which has gained growing attention in speech synthesis and image generation. In this paper, we introduce the novel notion of Ethically Sourced Code Generation (ES-CodeGen) to refer to managing all processes involved in code generation model development from data collection to post-deployment via ethical and sustainable practices. To build a taxonomy of ES-CodeGen, we perform a two-phase literature review where we reviewed 803 papers across various domains and specific to AI-based code generation. We identified 71 relevant papers with 10 initial dimensions of ES-CodeGen. To refine our dimensions and gain insights on consequences of ES-CodeGen, we surveyed 32 practitioners, which include six developers who submitted GitHub issues to opt-out from the Stack dataset (these impacted users have real-world experience of ethically sourced issues in code generation models). The results lead to 11 dimensions of ES-CodeGen with a new dimension on code quality as practitioners have noted its importance. We also identified consequences, artifacts, and stages relevant to ES-CodeGen. Our post-survey reflection showed that most practitioners tended to ignore society-related dimensions despite their importance. Most practitioners either agreed or strongly agreed that our survey help improve their understanding of ES-CodeGen. Our study calls for attention of various ethical issues towards ES-CodeGen.
• Software and its engineering → Software maintenance tools; • Human-centered computing → Empirical studies in collaborative and social computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- An Empirical Study on Software Bill of Materials: Where We Stand and the Road AheadBoming Xia, Tingting Bi, Zhenchang Xing, Qinghua Lu 等ICSE 2023 · 被引用 82 次
- Uncovering and Quantifying Social Biases in Code GenerationYan Liu, Xiaokang Chen, Yan Gao, Zhe Su 等NeurIPS 2023 · 被引用 47 次
- Using AI Assistants in Software Development: A Qualitative Study on Security Practices and ConcernsJan H. Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden 等CCS 2024 · 被引用 14 次
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu 等ASE 2025 · 被引用 2 次
- CodexLeaks: Privacy Leaks from Code Generation Language Models in GitHub CopilotLiang Niu, Muhammad Shujaat Mirza, Zayd Maradni, Christina PöpperUSENIX Security 2023
相关 Paper
- Do Users Write More Insecure Code with AI Assistants?Neil Perry, Megha Srivastava, Deepak Kumar, Dan BonehCCS 2023 · 被引用 150 次
- A User-centered Security Evaluation of CopilotOwura Asare, Meiyappan Nagappan, N. AsokanICSE 2024 · 被引用 11 次
- An Empirical Study of Code Clones from Commercial AI Code GeneratorsWeibin Wu, Haoxuan Hu, Zhaoji Fan, Yitong Qiao 等FSE 2025 · 被引用 2 次
- CodeIPPrompt: Intellectual Property Infringement Assessment of Code Language ModelsZhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang 等ICML 2023 · 被引用 42 次
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 被引用 6 次
