DAInfer: Inferring API Aliasing Specifications from Library Documentation via Neurosymbolic Optimization
Chengpeng Wang, Jipeng Zhang, Rongxin Wu, Charles Zhang
Abstract
Modern software systems heavily rely on various libraries, necessitating understanding API semantics in static analysis. However, summarizing API semantics remains challenging due to complex implementations or the unavailability of library code. This paper presents DAInfer, a novel approach for inferring API aliasing specifications from library documentation. Specifically, we employ Natural Language Processing (NLP) models to interpret informal semantic information provided by the documentation, which enables us to reduce the specification inference to an optimization problem. Furthermore, we propose a new technique called neurosymbolic optimization to efficiently solve the optimization problem, yielding the desired API aliasing specifications. We have implemented DAInfer as a tool and evaluated it upon Java classes from several popular libraries. The results indicate that DAInfer infers the API aliasing specifications with a precision of 79.78% and a recall of 82.29%, averagely consuming 5.35 seconds per class. These obtained aliasing specifications further facilitate alias analysis, revealing 80.05% more alias facts for API return values in 15 Java projects. Additionally, the tool supports taint analysis, identifying 85 more taint flows in 23 Android apps. These results demonstrate the practical value of DAInfer in library-aware static analysis.
CCS Concepts: • Software and its engineering → Software libraries and repositories; Automated static analysis; • Applied computing → Document analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c983368c-b9b7-4355-a278-d3922925be05Cited by top-tier papers3
- RFCAudit: AI Agent for Auditing Protocol Implementations Against RFC SpecificationsMingwei Zheng, Chengpeng Wang, Xuwei Liu, Jinyao Guo et al.ASE 2025 · 5 citations
- Validating Network Protocol Parsers with Traceable RFC Document InterpretationMingwei Zheng, Danning Xie, Qingkai Shi, Chengpeng Wang et al.ISSTA 2025 · 4 citations
- RepoAudit: An Autonomous LLM-Agent for Repository-Level Code AuditingJinyao Guo, Chengpeng Wang, Xiangzhe Xu, Zian Su et al.ICML 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Autoformalization with Large Language ModelsYuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe et al.NeurIPS 2022 · 364 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- HyperTree Proof Search for Neural Theorem ProvingGuillaume Lample, Timothée Lacroix, Marie-Anne Lachaux, Aurélien Rodriguez et al.NeurIPS 2022 · 271 citations
Related papers
- Broadening Horizons of Multilingual Static Analysis: Semantic Summary Extraction from C Code for JNI Program AnalysisSungho Lee, Hyogun Lee, Sukyoung RyuASE 2020 · 29 citations
- DocFlow: Extracting Taint Specifications from Software DocumentationMarcos Tileria, Jorge Blasco, Santanu Kumar DashICSE 2024 · 5 citations
- NESA: Relational Neuro-Symbolic Static Program AnalysisChengpeng Wang, Yifei Gao, Wuqi Zhang, Xuwei Liu et al.FSE 2026 · 1 citation
- NativeSummary: Summarizing Native Binary Code for Inter-language Static Analysis of Android AppsJikai Wang, Haoyu WangISSTA 2024 · 8 citations
- Interactive Cross-Language Pointer Analysis for Resolving Native Code in Java ProgramsChenxi Zhang, Yufei Liang, Tian Tan, Chang Xu et al.ICSE 2025 · 1 citation
