Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
James Jewitt, Gopi Krishnan Rajbahadur, Hao Li, Bram Adams, Ahmed E. Hassan
摘要
Permissive licenses like MIT, Apache-2.0, and BSD-3-Clause dominate open-source AI, signaling that artifacts like models, datasets, and code can be freely used, modified, and redistributed. However, these licenses carry mandatory requirements: include the full license text, provide a copyright notice, and preserve upstream attribution, that remain unverified at scale. Failure to meet these conditions can place reuse outside the scope of the license, effectively leaving AI artifacts under default copyright for those uses and exposing downstream users to litigation. We call this phenomenon "permissive washing": labeling AI artifacts as free to use, while omitting the legal documentation required to make that label actionable. To assess how widespread permissive washing is in the AI supply chain, we empirically audit 124,278 dataset → model → application supply chains, spanning 3,338 datasets, 6,664 models, and 28,516 applications across Hugging Face and GitHub. We find that an astonishing 96.5% of datasets and 95.8% of models lack the required license text, only 2.3% of datasets and 3.2% of models satisfy both license text and copyright requirements, and even when upstream artifacts provide complete licensing evidence, attribution rarely propagates downstream: only 27.59% of models preserve compliant dataset notices and only 5.75% of applications preserve compliant model notices (with just 6.38% preserving any linked upstream notice). Practitioners cannot assume permissive labels confer the rights they claim: license files and notices, not metadata, are the source of legal truth. To support future research, we release our full audit dataset and reproducible pipeline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- ModelGo: A Practical Tool for Machine Learning License AnalysisMoming Duan, Qinbin Li, Bingsheng HeWWW 2024 · 被引用 12 次
- Small Changes, Big Trouble: Demystifying and Parsing License Variants for Incompatibility Detection in the PyPI EcosystemWeiwei Xu, Hengzhi Ye, Kai Gao, Minghui ZhouICSE 2026 · 被引用 1 次
- Bridging the Data Provenance Gap Across Text, Speech, and VideoShayne Longpre, Nikhil Singh, Manuel Cherep, Kushagra Tiwary 等ICLR 2025
相关 Paper
- RAI2: Responsible Identity Audit Governing the Artificial IntelligenceTian Dong, Shaofeng Li, Guoxing Chen, Minhui Xue 等NDSS 2023
- Hidden Licensing Risks in the PTMware EcosystemBo Wang, Yueyang Chen, Jieke Shi, Minghui Li 等ISSTA 2026
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMsZhihao Wu, Gracia Gong, Qinglin Zhu, Yudong Chen 等ICML 2026
- IACW: Intent-Aware Controllable Watermarking for Scalable Authorial Intent AttributionHao Huang, Ruihua Zhou, JiaTang Luo, Yunpeng Li 等ICML 2026
- SoK: Dataset Copyright Auditing in Machine Learning SystemsLinkang Du, Xuanru Zhou, Min Chen, Chusong Zhang 等S&P 2025
