Finding Bugs Using Your Own Code: Detecting Functionally-similar yet Inconsistent Code
Mansour Ahmadi, Reza Mirzazade Farkhani, Ryan Williams, Long Lu
摘要
Probabilistic classification has shown success in detecting known types of software bugs. However, the works following this approach tend to require a large amount of specimens to train their models. We present a new machine learning-based bug detection technique that does not require any external code or samples for training. Instead, our technique learns from the very codebase on which the bug detection is performed, and therefore, obviates the need for the cumbersome task of gathering and cleansing training samples (e.g., buggy code of certain kinds). The key idea behind our technique is a novel two-step clustering process applied on a given codebase. This clustering process identifies code snippets in a project that are functionally-similar yet appear in inconsistent forms. Such inconsistencies are found to cause a wide range of bugs, anything from missing checks to unsafe type conversions. Unlike previous works, our technique is generic and not specific to one type of inconsistency or bug. We prototyped our technique and evaluated it using 5 popular open source software, including QEMU and OpenSSL. With a minimal amount of manual analysis on the inconsistencies detected by our tool, we discovered 22 new unique bugs, despite the fact that many of these programs are constantly undergoing bug scans and new bugs in them are believed to be rare.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Learning Graph-based Code Representations for Source-level Functional Similarity DetectionJiahao Liu, Jun Zeng, Xiang Wang, Zhenkai LiangICSE 2023 · 被引用 27 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Detecting Missed Security Operations Through Differential Checking of Object-based Similar PathsDinghao Liu, Qiushi Wu, Shouling Ji, Kangjie Lu 等CCS 2021 · 被引用 11 次
- One Bug, Hundreds Behind: LLMs for Large-Scale Bug DiscoveryQiushi Wu, Yue Xiao, Dhilung Kirat, Kevin Eykholt 等ICML 2026 · 被引用 5 次
- LineBreaker: Finding Token-Inconsistency Bugs with Large Language ModelsHongbo Chen, Yifan Zhang, Xing Han, Tianhao Mao 等ASE 2025 · 被引用 3 次
它引用的顶会 Paper7
- Neural Network-based Graph Embedding for Cross-Platform Binary Code Similarity DetectionXiaojun Xu, Chang Liu, Qian Feng, Heng Yin 等CCS 2017 · 被引用 682 次
- Scalable Graph-based Bug Search for Firmware ImagesQian Feng, Rundong Zhou, Chengcheng Xu, Yao Cheng 等CCS 2016 · 被引用 456 次
- A Large-Scale Empirical Study of Security PatchesFrank Li, Vern PaxsonCCS 2017 · 被引用 273 次
- APISan: Sanitizing API Usages through Semantic Cross-CheckingInsu Yun, Changwoo Min, Xujie Si, Yeongjin Jang 等USENIX Security 2016 · 被引用 107 次
- Detecting Missing-Check Bugs via Semantic- and Context-Aware Criticalness and Constraints InferencesKangjie Lu, Aditya Pakki, Qiushi WuUSENIX Security 2019 · 被引用 97 次
相关 Paper
- Self-Supervised Bug Detection and RepairMiltiadis Allamanis, Henry Jackson-Flux, Marc BrockschmidtNeurIPS 2021 · 被引用 145 次
- Non-Distinguishable Inconsistencies as a Deterministic Oracle for Detecting Security BugsQingyang Zhou, Qiushi Wu, Dinghao Liu, Shouling Ji 等CCS 2022 · 被引用 2 次
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 被引用 20 次
- How to Train Your Neural Bug Detector: Artificial vs Real BugsCedric Richter, Heike WehrheimASE 2023 · 被引用 5 次
- When to Say What: Learning to Find Condition-Message InconsistenciesIslem Bouzenia, Michael PradelICSE 2023 · 被引用 5 次
