Beware of the Unexpected: Bimodal Taint Analysis
Yiu Wai Chow, Max Schäfer, Michael Pradel
Abstract
Static analysis is a powerful tool for detecting security vulnerabilities and other programming problems. Global taint tracking, in particular, can spot vulnerabilities arising from complicated data flow across multiple functions. However, precisely identifying which flows are problematic is challenging, and sometimes depends on factors beyond the reach of pure program analysis, such as conventions and informal knowledge. For example, learning that a parameter name of an API function locale ends up in a file path is surprising and potentially problematic. In contrast, it would be completely unsurprising to find that a parameter command passed to an API function execaCommand is eventually interpreted as part of an operating-system command. This paper presents Fluffy, a bimodal taint analysis that combines static analysis, which reasons about data flow, with machine learning, which probabilistically determines which flows are potentially problematic. The key idea is to let machine learning models predict from natural language information involved in a taint flow, such as API names, whether the flow is expected or unexpected, and to inform developers only about the latter. We present a general framework and instantiate it with four learned models, which offer different trade-offs between the need to annotate training data and the accuracy of predictions. We implement Fluffy on top of the CodeQL analysis framework and apply it to 250K JavaScript projects. Evaluating on five common vulnerability types, we find that Fluffy achieves an F1 score of 0.85 or more on four of them across a variety of datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff4276e7-4367-47e3-ae1e-901166362c2bCited by top-tier papers5
- LLMDFA: Analyzing Dataflow in Code with Large Language ModelsChengpeng Wang, Wuqi Zhang, Zian Su, Xiangzhe Xu et al.NeurIPS 2024 · 51 citations
- SecBench.js: An Executable Security Benchmark Suite for Server-Side JavaScriptMasudul Hasan Masud Bhuiyan, Adithya Srinivas Parthasarathy, Nikos Vasilakis, Michael Pradel et al.ICSE 2023 · 20 citations
- Learning to Locate and Describe VulnerabilitiesJian Zhang, Shangqing Liu, Xu Wang, Tianlin Li et al.ASE 2023 · 8 citations
- CodeCureAgent: Automatic Classification and Repair of Static Analysis WarningsPascal Joos, Islem Bouzenia, Michael PradelFSE 2026
- Two Approaches to Fast Bytecode Frontend for Static AnalysisChenxi Li, Haoran Lin, Tian Tan, Yue LiOOPSLA 2025
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li et al.ISSTA 2020 · 325 citations
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 283 citations
- DLFix: context-based code transformation learning for automated program repairYi Li, Shaohua Wang, Tien N. NguyenICSE 2020 · 201 citations
- Repair Is Nearly Generation: Multilingual Program Repair with LLMsHarshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le et al.AAAI 2023 · 182 citations
Related papers
- Jasmine: Scale up JavaScript Static Security Analysis with Computation-based Semantic ExplanationFeng Xiao, Zhongfu Su, Guangliang Yang, Wenke LeeS&P 2024
- D-BUNDLR: Destructing JavaScript Bundles for Effective Static AnalysisWenyuan Xu, Alexi Turcotte, Cristian-Alexandru StaicuICSE 2026
- Reframing Paths as Logic: Semantic Segmentation for Vulnerability DetectionZong Cao, Yuqiang Sun, Zhengzi Xu, Kaixuan Li et al.OOPSLA 2026
- IRIS: LLM-Assisted Static Analysis for Detecting Security VulnerabilitiesZiyang Li, Saikat Dutta, Mayur NaikICLR 2025
- Towards a Lightweight, Hybrid Approach for Detecting DOM XSS Vulnerabilities with Machine LearningWilliam Melicher, Clement Fung, Lujo Bauer, Limin JiaWWW 2021 · 34 citations
