Is Call Graph Pruning Really Effective?: An Empirical Re-evaluation
Mohammad Rafieian, Vlad Birsan, Kunal Katiyar, Dylan Zhong, Shiyi Wei
Abstract
Recently proposed call graph pruning techniques, which use machine learning to predict and reduce false positives in static call graphs, have reported impressive results and demonstrated practicality for downstream analyses. However, as we research the methodology of how datasets are created and how evaluation is conducted in these research projects, we identify many risk factors that may make the reported results incomplete or even invalid. We further investigate these factors and provide empirical evidence to demonstrate they indeed result in misleading results. Motivated by these findings, we propose an empirical re-evaluation of existing call graph pruning techniques to assess their effectiveness. We first construct a new dataset of real-world programs, applying three complementary labeling approaches, resulting in 41,952 labels. To provide a full picture of these techniques, we utilize three popular call graph analysis frameworks (WALA, Doop, and OPAL) and seven tool configurations. Our results show that pruning is not as effective as reported in prior work, improving precision at the cost of significant coverage loss. We further train more general models across analysis tools and configurations, and explore the impact of data splitting and balancing. We find these models often match or outperform the existing configuration-specific ones, indicating their applicability in more general settings. Finally, we highlight the open challenges and potential solutions, including the need for better datasets as well as improved feature engineering.
• Software and its engineering → Automated static analysis; • Theory of computation → Program analysis; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2252a856-05d8-40b0-b268-105c4fa0a3c6Builds on15
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui et al.EMNLP 2023 · 339 citations
- Semi-supervised Log-based Anomaly Detection via Probabilistic Label EstimationLin Yang, Junjie Chen, Zan Wang, Weijing Wang et al.ICSE 2021 · 216 citations
- Making pointer analysis more precise by unleashing the power of selective context sensitivityTian Tan, Yue Li, Xiaoxing Ma, Chang Xu et al.OOPSLA 2021 · 39 citations
- Modular collaborative program analysis in OPALDominik Helm, Florian Kübler, Michael Reif, Michael Eichberg et al.FSE 2020 · 35 citations
- Learning to Reduce False Positives in Analytic Bug DetectorsAnant Kharkar, Roshanak Zilouchian Moghaddam, Matthew Jin, Xiaoyu Liu et al.ICSE 2022 · 33 citations
Related papers
- AutoPruner: transformer-based call graph pruningThanh Le-Cong, Hong Jin Kang, Truong Giang Nguyen, Stefanus Agus Haryono et al.FSE 2022 · 21 citations
- Striking a Balance: Pruning False-Positives from Static Call GraphsAkshay Utture, Shuyang Liu, Christian Gram Kalhauge, Jens PalsbergICSE 2022 · 18 citations
- Detecting False Alarms from Automatic Static Analysis Tools: How Far are We?Hong Jin Kang, Khai Loong Aw, David LoICSE 2022 · 42 citations
- Redefining Indirect Call Analysis with KallGraphGuoren Li, Manu Sridharan, Zhiyun QianS&P 2025
- An empirical study on the effectiveness of static C code analyzers for vulnerability detectionStephan Lipp, Sebastian Banescu, Alexander PretschnerISSTA 2022 · 99 citations
