Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects?
The Anh Nguyen, Triet Huynh Minh Le, M. Ali Babar
摘要
The rapid growth of Artificial Intelligence (AI) models and applications has led to an increasingly complex security landscape. Developers of AI projects must contend not only with traditional software supply chain issues but also with novel, AI-specific security threats. However, little is known about what security issues are commonly encountered and how they are resolved in practice. This gap hinders the development of effective security measures for each component of the AI supply chain. We bridge this gap by conducting an empirical investigation of developer-reported issues and solutions, based on discussions from Hugging Face and GitHub. To identify securityrelated discussions, we develop a pipeline that combines keyword matching with an optimal fine-tuned distilBERT classifier, which achieved the best performance in our extensive comparison of various deep learning and large language models. This pipeline produces a dataset of 312,868 security discussions, providing insights into the security reporting practices of AI applications and projects. We conduct a thematic analysis of 753 posts sampled from our dataset and uncover a fine-grained taxonomy of 32 security issues and 24 solutions across four themes: (1) System and Software, (2) External Tools and Ecosystem, (3) Model, and (4) Data. We reveal that many security issues arise from the complex dependencies and black-box nature of AI components. Notably, challenges related to Models and Data often lack concrete solutions. Our insights can offer evidence-based guidance for developers and researchers to address real-world security threats across the AI supply chain.
• Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finding A Needle in a Haystack: Automated Mining of Silent Vulnerability FixesJiayuan Zhou, Michael Pacheco, Zhiyuan Wan, Xin Xia 等ASE 2021 · 被引用 84 次
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer 等ICSE 2023 · 被引用 62 次
相关 Paper
- Your Space is My Zone: Demystifying the Security Risks of AI-Powered Applications on Pre-Trained Model HubsYacong Gu, Lingyun Ying, Zidong Zhang, Yingyuan Pu 等CCS 2026 · 被引用 1 次
- A First Look at Model Supply Chain: From the Risk PerspectiveZiqian Chen, Zekai Chen, Susheng Wu, Bihuan Chen 等ICSE 2026 · 被引用 1 次
- Towards More Practical Threat Models in Artificial Intelligence SecurityKathrin Grosse, Lukas Bieringer, Tarek R. Besold, Alexandre AlahiUSENIX Security 2024 · 被引用 27 次
- A Large-Scale Empirical Study of Secret Key Leakage in Hugging Face SpacesShaoxuan Yun, Yuchao Zhang, Zhikun Shi, Liu Wang 等ICSE 2026
- Using AI Assistants in Software Development: A Qualitative Study on Security Practices and ConcernsJan H. Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden 等CCS 2024 · 被引用 14 次
