The Product Beyond the Model - An Empirical Study of Repositories of Open-Source ML Products
Nadia Nahar, Haoran Zhang, Grace A. Lewis, Shurui Zhou, Christian Kästner
摘要
Machine learning (ML) components are increasingly incorporated into software products for end-users, but developers face challenges in transitioning from ML prototypes to products. Academics have limited access to the source of commercial ML products, hindering research progress to address these challenges. In this study, first and foremost, we contribute a dataset of 262 open-source ML products for end users (not just models), identified among more than half a million ML-related projects on GitHub. Then, we qualitatively and quantitatively analyze 30 open-source ML products to answer six broad research questions about development practices and system architecture. We find that the majority of the ML products in our sample represent more startup-style development than reported in past interview studies. We report 21 findings, including limited involvement of data scientists in many open-source ML products, unusually low modularity between ML and non-ML code, diverse architectural choices on incorporating models into products, and limited prevalence of industry best practices such as model testing, pipeline automation, and monitoring. Additionally, we discuss seven implications of this study on research, development, and education, including the need for tools to assist teams without data scientists, education opportunities, and open-source-specific research for privacy-preserving telemetry.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Large-Scale Study of Model Integration in ML-Enabled Software SystemsYorick Sens, Henriette Knopp, Sven Peldszus, Thorsten BergerICSE 2025 · 被引用 3 次
- Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph ExecutionRaffi Khatchadourian, Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia 等ASE 2025 · 被引用 1 次
它引用的顶会 Paper14
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for Shifting Organizational PracticesBogdana Rakova, Jingying Yang, Henriette Cramer, Rumman ChowdhuryCSCW 2021 · 被引用 326 次
- How AI Developers Overcome Communication Challenges in a Multidisciplinary Team: A Case StudyDavid Piorkowski, Soya Park, April Yi Wang, Dakuo Wang 等CSCW 2021 · 被引用 142 次
- Collaboration Challenges in Building ML-Enabled Systems: Communication, Documentation, Engineering, and ProcessNadia Nahar, Shurui Zhou, Grace A. Lewis, Christian KästnerICSE 2022 · 被引用 122 次
- Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training DataKaren L. BoydCSCW 2021 · 被引用 68 次
相关 Paper
- Analyzing Collaborative Challenges and Needs of UX Practitioners when Designing with AI/MLMeena Devii Muralikumar, David W. McDonaldCSCW 2024 · 被引用 6 次
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang 等FSE 2020 · 被引用 121 次
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire 等ICSE 2021 · 被引用 37 次
- Towards Observability for Production Machine Learning Pipelines [Vision]Shreya Shankar, Aditya G. ParameswaranVLDB 2022 · 被引用 21 次
- 23 shades of self-admitted technical debt: an empirical study on machine learning softwareDavid O'Brien, Sumon Biswas, Sayem Imtiaz, Rabe Abdalkareem 等FSE 2022 · 被引用 37 次
