Identifying Multi-parameter Constraint Errors in Python Data Science Library API Documentation
Xiufeng Xu, Fuman Xie, Chenguang Zhu, Guangdong Bai, Sarfraz Khurshid, Yi Li
摘要
Modern AI- and Data-intensive software systems rely heavily on data science and machine learning libraries that provide essential algorithmic implementations and computational frameworks. These libraries expose complex APIs whose correct usage has to follow constraints among multiple interdependent parameters. Developers using these APIs are expected to learn about the constraints through the provided documentation and any discrepancy may lead to unexpected behaviors. However, maintaining correct and consistent multi-parameter constraints in API documentation remains a significant challenge for API compatibility and reliability. To address this challenge, we propose MPChecker for detecting inconsistencies between code and documentation, specifically focusing on multi-parameter constraints. MPChecker identifies these constraints at the code level by exploring execution paths through symbolic execution and further extracts corresponding constraints from documentation using large language models (LLMs). We propose a customized fuzzy constraint logic to reconcile the unpredictability of LLM outputs and detect logical inconsistencies between the code and documentation constraints. We collected and constructed two datasets from four popular data science libraries and evaluated MPChecker on them. Our tool identified 117 of 126 inconsistent constraints, achieving a recall of 92.8% and demonstrating its effectiveness at detecting inconsistency issues. We further reported 14 detected inconsistency issues to the library developers, who have confirmed 11 issues at the time of writing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu 等ICSE 2024 · 被引用 264 次
- LLM Maybe LongLM: SelfExtend LLM Context Window Without TuningHongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang 等ICML 2024 · 被引用 167 次
- Software documentation: the practitioners' perspectiveEmad Aghajani, Csaba Nagy, Mario Linares-Vásquez, Laura Moreno 等ICSE 2020 · 被引用 112 次
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 被引用 105 次
相关 Paper
- DocTer: documentation-guided fuzzing for testing deep learning API functionsDanning Xie, Yitong Li, Mijung Kim, Hung Viet Pham 等ISSTA 2022 · 被引用 72 次
- Can LLMs Implicitly Learn Numeric Parameter Constraints in Data Science APIs?Yinlin Deng, Chunqiu Steven Xia, Zhezhen Cao, Meiziniu Li 等NeurIPS 2024 · 被引用 4 次
- CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test GenerationTobias Kiecker, Jan Arne Sparka, Martin Reuter, Albert Ziegler 等FSE 2026 · 被引用 1 次
- The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse DetectionYi Yang, Jinghua Liu, Kai Chen, Miaoqian LinNDSS 2025
- Oracle-Guided Program Selection from Large Language ModelsZhiyu Fan, Haifeng Ruan, Sergey Mechtaev, Abhik RoychoudhuryISSTA 2024 · 被引用 4 次
