Demystifying and Detecting Misuses of Deep Learning APIs
Moshi Wei, Nima Shiri Harzevili, Yuekai Huang, Jinqiu Yang, Junjie Wang, Song Wang
摘要
Deep Learning (DL) libraries have significantly impacted various domains in computer science over the last decade. However, developers often face challenges when using the DL APIs, as the development paradigm of DL applications differs greatly from traditional software development. Existing studies on API misuse mainly focus on traditional software, leaving a gap in understanding API misuse within DL APIs. To address this gap, we present the first comprehensive study of DL API misuse in TensorFlow and PyTorch. Specifically, we first collect a dataset of 4,224 commits from the top 200 most-starred projects using these two libraries and manually identified 891 API misuses. We then investigate the characteristics of these misuses from three perspectives, i.e., types, root causes, and symptoms. We have also conducted an evaluation to assess the effectiveness of the current state-of-the-art API misuse detector on our 891 confirmed API misuses. Our results confirmed that the stateof-the-art (SOTA) API misuse detector is ineffective in detecting DL API misuses. To address the limitations of existing API misuse detection for DL APIs, we propose LLMAPIDet, which leverages Large Language Models (LLMs) for DL API misuse detection and repair. We build LLMAPIDet by prompt-tuning a chain of ChatGPT prompts on 600 out of 891 confirmed API misuses and reserve the rest 291 API misuses as the testing dataset. Our evaluation shows that LLMAPIDet can detect 48 out of the 291 DL API misuses while none of them can be detected by the existing API misuse detector. We further evaluate LLMAPIDet on the latest versions of 10 GitHub projects. The evaluation shows that LLMAPIDet can identify 119 previously unknown API misuses and successfully fix 46 of them. CCS CONCEPTS • Software and its engineering → Software evolution; Software libraries and repositories; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Are LLMs Correctly Integrated into Software Systems?Yuchen Shao, Yuheng Huang, Jiawei Shen, Lei Ma 等ICSE 2025 · 被引用 4 次
- LineBreaker: Finding Token-Inconsistency Bugs with Large Language ModelsHongbo Chen, Yifan Zhang, Xing Han, Tianhao Mao 等ASE 2025 · 被引用 3 次
- Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph ExecutionRaffi Khatchadourian, Tatiana Castro Vélez, Mehdi Bagherzadeh, Nan Jia 等ASE 2025 · 被引用 1 次
- My Model is Malware to You: Transforming AI Models into Malware by Abusing TensorFlow APIsRuofan Zhu, Ganhao Chen, Wenbo Shen, Xiaofei Xie 等S&P 2025
- Cryptbara: Dependency-Guided Detection of Python Cryptographic API MisusesSeogyeong Cho, Seungeun Yu, Seunghoon WooASE 2025
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Taxonomy of real faults in deep learning systemsNargiz Humbatova, Gunel Jahangirova, Gabriele Bavota, Vincenzo Riccio 等ICSE 2020 · 被引用 281 次
- Repository-Level Prompt Generation for Large Language Models of CodeDisha Shrivastava, Hugo Larochelle, Daniel TarlowICML 2023 · 被引用 184 次
- Coder Reviewer Reranking for Code GenerationTianyi Zhang, Tao Yu, Tatsunori Hashimoto, Mike Lewis 等ICML 2023 · 被引用 125 次
相关 Paper
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu 等ICSE 2025 · 被引用 6 次
- The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse DetectionYi Yang, Jinghua Liu, Kai Chen, Miaoqian LinNDSS 2025
- LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-Based Code CompletionChong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng 等ICSE 2025 · 被引用 3 次
- The Seeds of the Future Sprout from History: Fuzzing for Unveiling Vulnerabilities in Prospective Deep-Learning LibrariesZhiyuan Li, Jingzheng Wu, Xiang Ling, Tianyue Luo 等ICSE 2025 · 被引用 4 次
- Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMsYifan Xia, Zichen Xie, Peiyu Liu, Kangjie Lu 等ISSTA 2025 · 被引用 2 次
