Combining Structured Static Code Information and Dynamic Symbolic Traces for Software Vulnerability Prediction
Huanting Wang, Zhanyong Tang, Shin Hwei Tan, Jie Wang, Yuzhe Liu, Hejun Fang, Chunwei Xia, Zheng Wang
摘要
Deep learning (DL) has emerged as a viable means for identifying software bugs and vulnerabilities. The success of DL relies on having a suitable representation of the problem domain. However, existing DL-based solutions for learning program representations have limitations -they either cannot capture the deep, precise program semantics or suffer from poor scalability. We present Concoction, the first DL system to learn program presentations by combining static source code information and dynamic program execution traces. Concoction employs unsupervised active learning techniques to determine a subset of important paths to collect dynamic symbolic execution traces. By implementing a focused symbolic execution solution, Concoction brings the benefits of static and dynamic code features while reducing the expensive symbolic execution overhead. We integrate Concoction with fuzzing techniques to detect function-level code vulnerabilities in C programs from 20 open-source projects. In 200 hours of automated concurrent test runs, Concoction has successfully uncovered vulnerabilities in all tested projects, identifying 54 unique vulnerabilities and yielding 37 new, unique CVE IDs. Concoction also significantly outperforms 16 prior methods by providing higher accuracy and lower false positive rates. CCS CONCEPTS • Security and privacy → Software security engineering; • Computing methodologies → Artificial intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Rethinking Code Similarity for Automated Algorithm Design with LLMsRui Zhang, Zhichao LuICLR 2026 · 被引用 2 次
- Similar but Patched Code Considered Harmful: The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove ThemZixuan Tan, Jiayuan Zhou, Xing Hu, Shengyi Pan 等ICSE 2025 · 被引用 1 次
- NCFuzz: Configuration-Guided Network Service FuzzingXuesong Bai, Hengkai Ye, Shenghan Zheng, Fenglu Zhang 等ISSTA 2026
它引用的顶会 Paper26
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
相关 Paper
- VulCNN: An Image-inspired Scalable Vulnerability Detection SystemYueming Wu, Deqing Zou, Shihan Dou, Wei Yang 等ICSE 2022 · 被引用 141 次
- Lightweight Concolic Testing via Path-Condition Synthesis for Deep Learning LibrariesSehoon Kim, Yonghyeon Kim, Dahyeon Park, Yuseok Jeon 等ICSE 2025 · 被引用 5 次
- Path-sensitive code embedding via contrastive learning for software vulnerability detectionXiao Cheng, Guanqin Zhang, Haoyu Wang, Yulei SuiISSTA 2022 · 被引用 98 次
- A Generative and Mutational Approach for Synthesizing Bug-Exposing Test Cases to Guide Compiler FuzzingGuixin Ye, Tianmin Hu, Zhanyong Tang, Zhenye Fan 等FSE 2023 · 被引用 14 次
- Learning to Fuzz from Symbolic Execution with Application to Smart ContractsJingxuan He, Mislav Balunovic, Nodar Ambroladze, Petar Tsankov 等CCS 2019 · 被引用 288 次
