Automating API Documentation from Crowdsourced Knowledge
Bonan Kou, Zijie Zhou, Muhao Chen, Tianyi Zhang
Abstract
API documentation is crucial for developers to learn and use APIs. However, it is known that many official API documents are obsolete and incomplete. To address this challenge, we propose a new approach called AutoDoc that generates API documents with API knowledge extracted from online discussions on Stack Overflow (SO). AutoDoc leverages a fine-tuned dense retrieval model to identify seven types of API knowledge from SO posts. Then, it uses GPT-4o to summarize the API knowledge in these posts into concise text. Meanwhile, we designed two specific components to handle LLM hallucination and redundancy in generated content.
We evaluated AutoDoc against five comparison baselines on 48 APIs of different popularity levels. Our results indicate that the API documents generated by AutoDoc are up to 77.7% more accurate, 9.5% less duplicated, and contain 34.4% knowledge uncovered by the official documents. We also measured the sensitivity of AutoDoc to the choice of different LLMs. We found that while larger LLMs produce higher-quality API documents, AutoDoc enables smaller open-source models (e.g., Mistral-7B-v0.3) to achieve comparable results. Finally, we conducted a user study to evaluate the usefulness of the API documents generated by AutoDoc. All participants found API documents generated by AutoDoc to be more comprehensive, concise, and helpful than the comparison baselines. This highlights the feasibility of utilizing LLMs for API documentation with careful design to counter LLM hallucination and information redundancy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2bef557-13a1-42de-8e07-dcaecf9e90f6Builds on5
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- CollabCoder: A Lower-barrier, Rigorous Workflow for Inductive Collaborative Qualitative Analysis with Large Language ModelsJie Gao, Yuchen Guo, Gionnieve Lim, Tianqin Zhang et al.CHI 2024 · 63 citations
- Demystify official API usage directives with crowdsourced API misuse scenarios, erroneous code examples and patchesXiaoxue Ren, Jiamou Sun, Zhenchang Xing, Xin Xia et al.ICSE 2020 · 28 citations
- Automatic Solution Summarization for Crash BugsHaoye Wang, Xin Xia, David Lo, John C. Grundy et al.ICSE 2021 · 14 citations
- Automated Summarization of Stack Overflow PostsBonan Kou, Muhao Chen, Tianyi ZhangICSE 2023 · 7 citations
Related papers
- Natural Language-Focused Software Engineering via Code-Documentation EquivalenceAryaz Eghbali, Zhongxin Liu, Michael PradelFSE 2026
- Speculate: Generating REST API Specifications using LLMsKrishanu Singh, Kushagra Karar, Abhilash Jindal, Guowei YangFSE 2026
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- Escaping Whack-a-Mole: Optimizing Documentation as Repo-Specific Playbooks for Coding AgentsYutong Cheng, Haifeng Chen, Wenchao Yu, Xujiang Zhao et al.ICML 2026
- API Pack: A Massive Multi-Programming Language Dataset for API Call GenerationZhen Guo, Adriana Meza Soria, Wei Sun, Yikang Shen et al.ICLR 2025
