RTFM! Automatic Assumption Discovery and Verification Derivation from Library Document for API Misuse Detection
Tao Lv, Ruishi Li, Yi Yang, Kai Chen, Xiaojing Liao, XiaoFeng Wang, Peiwei Hu, Luyi Xing
Abstract
To use library APIs, a developer is supposed to follow guidance and respect some constraints, which we call integration assumptions (IAs). Violations of these assumptions can have serious consequences, introducing security-critical flaws such as use-after-free, NULL-dereference, and authentication errors. Analyzing a program for compliance with IAs involves significant effort and needs to be automated. A promising direction is to automatically recover IAs from a library document using Natural Language Processing (NLP) and then verify their consistency with the ways APIs are used in a program through code analysis. However, a practical solution along this line needs to overcome several key challenges, particularly the discovery of IAs from loosely formatted documents and interpretation of their informal descriptions to identify complicated constraints (e.g., data-/control-flow relations between different APIs). In this paper, we present a new technique for automated assumption discovery and verification derivation from library documents. Our approach, called Advance, utilizes a suite of innovations to address those challenges. More specifically, we leverage the observation that IAs tend to express a strong sentiment in emphasizing the importance of a constraint, particularly those security-critical, and utilize a new sentiment analysis model to accurately recover them from loosely formatted documents. These IAs are further processed to identify hidden references to APIs and parameters, through an embedding model, to identify the information-flow relations expected to be followed. Then our approach runs frequent subtree mining to discover the grammatical units in IA sentences that tend to indicate some categories of constraints that could have security implications. These components are mapped to verification code snippets organized in line with the IA sentence's grammatical structure, and can be assembled into verification code executed through CodeQL to discover misuses inside a program. We implemented this design and evaluated it on 5 popular libraries (OpenSSL, SQLite, libpcap, libdbus and libxml2) and 39 real-world applications. Our analysis discovered 193 API misuses, including 139 flaws never reported before.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- DocTer: documentation-guided fuzzing for testing deep learning API functionsDanning Xie, Yitong Li, Mijung Kim, Hung Viet Pham et al.ISSTA 2022 · 72 citations
- Prompt Fuzzing for Fuzz Driver GenerationYunlong Lyu, Yuxuan Xie, Peng Chen, Hao ChenCCS 2024 · 21 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Breaking the Specification: PDF CertificationSimon Rohlmann, Vladislav Mladenov, Christian Mainka, Jörg SchwenkS&P 2021 · 16 citations
- Cerberus: Query-driven Scalable Vulnerability Detection in OAuth Service Provider ImplementationsTamjid Al Rahat, Yu Feng, Yuan TianCCS 2022 · 12 citations
Builds on4
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- Automated Analysis of Privacy Requirements for Mobile AppsSebastian Zimmeck, Ziqi Wang, Lieyong Zou, Roger Iyengar et al.NDSS 2017 · 255 citations
- APISan: Sanitizing API Usages through Semantic Cross-CheckingInsu Yun, Changwoo Min, Xujie Si, Yeongjin Jang et al.USENIX Security 2016 · 107 citations
- Devils in the Guidance: Predicting Logic Vulnerabilities in Payment Syndication Services through Automated Documentation AnalysisYi Chen, Luyi Xing, Yue Qin, Xiaojing Liao et al.USENIX Security 2019 · 33 citations
Related papers
- APICAD: Augmenting API Misuse Detection through Specifications from Code and DocumentsXiaoke Wang, Lei ZhaoICSE 2023 · 7 citations
- API-Misuse Detection Driven by Fine-Grained API-Constraint Knowledge GraphXiaoxue Ren, Xinyuan Ye, Zhenchang Xing, Xin Xia et al.ASE 2020 · 62 citations
- An empirical study on API parameter rulesHao Zhong, Na Meng, Zexuan Li, Li JiaICSE 2020 · 16 citations
- Generating API Parameter Security Rules with LLM for API Misuse DetectionJinghua Liu, Yi Yang, Kai Chen, Miaoqian LinNDSS 2025
- SFA-Miner: Mining Path-Sensitive API Usage Patterns Via Symbolic Finite AutomataJiasheng Jiang, Mingwei Zheng, Qingkai Shi, Xiangyu ZhangS&P 2026 · 1 citation
