CodeIPPrompt: Intellectual Property Infringement Assessment of Code Language Models
Zhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang, Yevgeniy Vorobeychik, Chaowei Xiao
摘要
Recent advances in large language models (LMs) have facilitated their ability to synthesize programming code. However, they have also raised concerns about intellectual property (IP) rights violations. Despite the significance of this issue, it has been relatively less explored. In this paper, we aim to bridge the gap by presenting CODEIPPROMPT, a platform for automatic evaluation of the extent to which code language models may reproduce licensed programs. It comprises two key components: prompts constructed from a licensed code database to elicit LMs to generate IP-violating code, and a measurement tool to evaluate the extent of IP violation of code LMs. We conducted an extensive evaluation of existing open-source code LMs and commercial products, and revealed the prevalence of IP violations in all these models. We further identified that the root cause is the substantial proportion of training corpus subject to restrictive licenses, resulting from both intentional inclusion and inconsistent license practice in the real world. To address this issue, we also explored potential mitigation strategies, including fine-tuning and dynamic token filtering.
Our study provides a testbed for evaluating the IP violation issues of the existing code generation platforms and stresses the need for a better mitigation strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron 等USENIX Security 2024 · 被引用 103 次
- Licoeval: Evaluating LLMs on License Compliance in Code GenerationWeiwei Xu, Kai Gao, Hao He, Minghui ZhouICSE 2025 · 被引用 6 次
- Beyond Raw Detection Scores: Markov-Informed Calibration for Boosting Machine-Generated Text DetectionChenwang Wu, Yiu-ming Cheung, Shuhai Zhang, Bo Han 等ICLR 2026 · 被引用 2 次
- An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far are We?Hyunjae Suh, Mahan Tafreshipour, Jiawei Li, Adithya Bhattiprolu 等ICSE 2025 · 被引用 2 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Copyright-Bench: Agentic Evaluation of Copyright Law ComplianceZheng Hui, Doni Bloomfield, Noam KoltICML 2026 · 被引用 1 次
- RMCBench: Benchmarking Large Language Models' Resistance to Malicious CodeJiachi Chen, Qingyuan Zhong, Yanlin Wang, Kaiwen Ning 等ASE 2024 · 被引用 6 次
- A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated CodeAlejandro Velasco, Daniel Rodriguez-Cardenas, Dipin Khati, David N. Palacio 等ICSE 2026
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 被引用 4 次
