Beware of the Unexpected: Bimodal Taint Analysis
Yiu Wai Chow, Max Schäfer, Michael Pradel
摘要
Static analysis is a powerful tool for detecting security vulnerabilities and other programming problems. Global taint tracking, in particular, can spot vulnerabilities arising from complicated data flow across multiple functions. However, precisely identifying which flows are problematic is challenging, and sometimes depends on factors beyond the reach of pure program analysis, such as conventions and informal knowledge. For example, learning that a parameter name of an API function locale ends up in a file path is surprising and potentially problematic. In contrast, it would be completely unsurprising to find that a parameter command passed to an API function execaCommand is eventually interpreted as part of an operating-system command. This paper presents Fluffy, a bimodal taint analysis that combines static analysis, which reasons about data flow, with machine learning, which probabilistically determines which flows are potentially problematic. The key idea is to let machine learning models predict from natural language information involved in a taint flow, such as API names, whether the flow is expected or unexpected, and to inform developers only about the latter. We present a general framework and instantiate it with four learned models, which offer different trade-offs between the need to annotate training data and the accuracy of predictions. We implement Fluffy on top of the CodeQL analysis framework and apply it to 250K JavaScript projects. Evaluating on five common vulnerability types, we find that Fluffy achieves an F1 score of 0.85 or more on four of them across a variety of datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LLMDFA: Analyzing Dataflow in Code with Large Language ModelsChengpeng Wang, Wuqi Zhang, Zian Su, Xiangzhe Xu 等NeurIPS 2024 · 被引用 51 次
- SecBench.js: An Executable Security Benchmark Suite for Server-Side JavaScriptMasudul Hasan Masud Bhuiyan, Adithya Srinivas Parthasarathy, Nikos Vasilakis, Michael Pradel 等ICSE 2023 · 被引用 20 次
- Learning to Locate and Describe VulnerabilitiesJian Zhang, Shangqing Liu, Xu Wang, Tianlin Li 等ASE 2023 · 被引用 8 次
- CodeCureAgent: Automatic Classification and Repair of Static Analysis WarningsPascal Joos, Islem Bouzenia, Michael PradelFSE 2026
- Two Approaches to Fast Bytecode Frontend for Static AnalysisChenxi Li, Haoran Lin, Tian Tan, Yue LiOOPSLA 2025
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CoCoNuT: combining context-aware neural translation models using ensemble for program repairThibaud Lutellier, Hung Viet Pham, Lawrence Pang, Yitong Li 等ISSTA 2020 · 被引用 325 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- DLFix: context-based code transformation learning for automated program repairYi Li, Shaohua Wang, Tien N. NguyenICSE 2020 · 被引用 201 次
- Repair Is Nearly Generation: Multilingual Program Repair with LLMsHarshit Joshi, José Pablo Cambronero Sánchez, Sumit Gulwani, Vu Le 等AAAI 2023 · 被引用 182 次
相关 Paper
- Jasmine: Scale up JavaScript Static Security Analysis with Computation-based Semantic ExplanationFeng Xiao, Zhongfu Su, Guangliang Yang, Wenke LeeS&P 2024
- D-BUNDLR: Destructing JavaScript Bundles for Effective Static AnalysisWenyuan Xu, Alexi Turcotte, Cristian-Alexandru StaicuICSE 2026
- Reframing Paths as Logic: Semantic Segmentation for Vulnerability DetectionZong Cao, Yuqiang Sun, Zhengzi Xu, Kaixuan Li 等OOPSLA 2026
- IRIS: LLM-Assisted Static Analysis for Detecting Security VulnerabilitiesZiyang Li, Saikat Dutta, Mayur NaikICLR 2025
- Towards a Lightweight, Hybrid Approach for Detecting DOM XSS Vulnerabilities with Machine LearningWilliam Melicher, Clement Fung, Lujo Bauer, Limin JiaWWW 2021 · 被引用 34 次
