Licoeval: Evaluating LLMs on License Compliance in Code Generation
Weiwei Xu, Kai Gao, Hao He, Minghui Zhou
摘要
Recent advances in Large Language Models (LLMs) have revolutionized code generation, leading to widespread adoption of AI coding tools by developers. However, LLMs can generate license-protected code without providing the necessary license information, leading to potential intellectual property violations during software production. This paper addresses the critical, yet underexplored, issue of license compliance in LLM-generated code by establishing a benchmark to evaluate the ability of LLMs to provide accurate license information for their generated code. To establish this benchmark, we conduct an empirical study to identify a reasonable standard for “striking similarity” that excludes the possibility of independent creation, indicating a copy relationship between the LLM output and certain opensource code. Based on this standard, we propose LiCoEval, to evaluate the license compliance capabilities of LLMs, i.e., the ability to provide accurate license or copyright information when they generate code with striking similarity to already existing copyrighted code. Using LiCoEval, we evaluate 14 popular LLMs, finding that even top-performing LLMs produce a non-negligible proportion (0.88 % to 2.01 %) of code strikingly similar to existing open-source implementations. Notably, most LLMs fail to provide accurate license information, particularly for code under copyleft licenses. These findings underscore the urgent need to enhance LLM compliance capabilities in code generation tasks. Our study provides a foundation for future research and development to improve license compliance in AIassisted software development, contributing to both the protection of open-source software copyrights and the mitigation of legal risks for LLM users.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- An Eye for AI: Eye-Tracking the Micro-Interruptions of GenAI Code SuggestionsTarek Alakmeh, Sarah D’Angelo, Thomas FritzICSE 2026 · 被引用 1 次
- Small Changes, Big Trouble: Demystifying and Parsing License Variants for Incompatibility Detection in the PyPI EcosystemWeiwei Xu, Hengzhi Ye, Kai Gao, Minghui ZhouICSE 2026 · 被引用 1 次
- How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?Hua Yang, Alejandro Velasco, Thanh Le-Cong, Md Nazmul Haque 等ICSE 2026
- Large Language Model Unlearning for Source CodeXue Jiang, Yihong Dong, Huangzhao Zhang, Tangxinyu Wang 等AAAI 2026
它引用的顶会 Paper20
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- Counterfactual Memorization in Neural Language ModelsChiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski 等NeurIPS 2023 · 被引用 184 次
相关 Paper
- CodeIPPrompt: Intellectual Property Infringement Assessment of Code Language ModelsZhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang 等ICML 2023 · 被引用 42 次
- Can Large Language Models Write Parallel Code?Daniel Nichols, Joshua Hoke Davis, Zhaojun Xie, Arjun Rajaram 等HPDC 2024 · 被引用 30 次
- Copyright-Bench: Agentic Evaluation of Copyright Law ComplianceZheng Hui, Doni Bloomfield, Noam KoltICML 2026 · 被引用 1 次
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far are We?Hyunjae Suh, Mahan Tafreshipour, Jiawei Li, Adithya Bhattiprolu 等ICSE 2025 · 被引用 2 次
- RMCBench: Benchmarking Large Language Models' Resistance to Malicious CodeJiachi Chen, Qingyuan Zhong, Yanlin Wang, Kaiwen Ning 等ASE 2024 · 被引用 6 次
