Exploring Static Taint Analysis in LLMs: A Dynamic Benchmarking Framework for Measurement and Enhancement
Haoran Zhao, Lei Zhang, Keke Lian, Fute Sun, Bofei Chen, Yongheng Liu, Zhiyu Wu, Yuan Zhang, Min Yang
Abstract
LLMs offer a promising avenue to overcome the limitations of traditional taint analysis techniques, with a growing number of studies leveraging LLMs for taint analysis and its downstream applications. However, these studies lack a systematic understanding of LLMs’ taint analysis capabilities, limiting their transferability and reliability. To bridge this gap and better apply LLMs to static taint analysis, we aim to comprehensively measure and understand LLMs’ taint analysis capabilities.Using existing benchmarks is a straightforward approach, but they are unsuitable due to issues such as training data leakage, not accounting for LLMs’ features, and improper assessment criteria. Manually constructing new benchmarks is not only labor-intensive but also struggles to remain effective as LLMs evolve. To address these, we propose LLMCapLens, a dynamic benchmark generation framework to systematically measure and enhance LLMs’ capabilities. LLMCapLens models influencing factors of LLMs’ taint analysis capabilities, employing a Basic Unit-Based generation method and a lightweight dynamic taint analysis-based verification method to implement the automated generation of targeted benchmarks, ensuring both diversity and correctness. Furthermore, LLMCapLens proposes a measurement-driven, training-free, model-specific enhancement approach.We apply LLMCapLens to 10 mainstream LLMs, revealing how they perform under various influencing factors and identifying unique characteristics, such as the underlying error causes for each model. Notably, our enhancement approach significantly improves LLM performance—GPT-4 Turbo, for instance, achieved improvements across 16 out of 19 factors, with an average True Negative Rate increase of 21.29%. Finally, we validate the real-world impact of our method by applying enhanced LLMs to vulnerability detection, demonstrating a substantial improvement over prior approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu et al.ICSE 2024 · 264 citations
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and BenchmarksSaad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce et al.S&P 2024 · 167 citations
Related papers
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 5 citations
- TaintP2X: Detecting Taint-Style Prompt-to-Anything Injection Vulnerabilities in LLM-Integrated ApplicationsJunjie He, Shenao Wang, Yanjie Zhao, Xinyi Hou et al.ICSE 2026
- Exploring and Improving Real-World Vulnerability Data Generation via Prompting Large Language ModelsGuangbei Yi, Yu Nong, Minzhang Li, Haipeng CaiICSE 2026
- MGTBench: Benchmarking Machine-Generated Text DetectionXinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes et al.CCS 2024 · 30 citations
- Quantifying Frontier LLM Capabilities for Container Sandbox EscapeRahul Marchand, Art Cathain, Jerome Wynne, Philippos Giavridis et al.ICML 2026 · 9 citations
