Igor: Crash Deduplication Through Root-Cause Clustering
Zhiyuan Jiang, Xiyue Jiang, Ahmad Hazimeh, Chaojing Tang, Chao Zhang, Mathias Payer
摘要
Fuzzing has emerged as the most effective bug-finding technique. The output of a fuzzer is a set of proof-of-concept (PoC) test cases for all observed "unique'' crashes. It costs developers substantial efforts to analyze each crashing test case. This, mostly manual, process has lead to the number of reported crashes out-pacing the number of bug fixes. Automatic crash deduplication techniques, which mostly rely on coverage profiles and stack hashes, are supposed to alleviate these pressures. However, these techniques both inflate actual bug counts and falsely conflate unrelated bugs. This hinders, rather than helps, developers, and calls for more accurate techniques. The highly-stochastic nature of fuzzing means that PoCs commonly exercise many program behaviors that are orthogonal to the crash's underlying root cause. This diversity in program behaviors manifests as a diversity in crashes, contributing to bug-count inflation and conflation. Based on this insight, we develop Igor, an automated dual-phase crash deduplication technique. By minimizing each PoC's execution trace, we obtain pruned test cases that exercise the critical behavior necessary for triggering a bug. Then, we use a graph similarity comparison to cluster crashes based on the control-flow graph of the minimized execution traces, with each cluster mapping back to a single, unique root cause. We evaluate Igor against 39 bugs resulting from 254,000 PoCs, distributed over 10 programs. Our results show that Igor accurately groups these crashes into 48 uniquely identifiable clusters, while other state-of-the-art methods yield bug counts at least one order of magnitude larger.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- JIT-Picking: Differential Fuzzing of JavaScript EnginesLukas Bernhard, Tobias Scharnowski, Moritz Schloegel, Tim Blazytko 等CCS 2022 · 被引用 42 次
- Everything is Good for Something: Counterexample-Guided Directed Fuzzing via Likely Invariant InferenceHeqing Huang, Anshunkang Zhou, Mathias Payer, Charles ZhangS&P 2024 · 被引用 15 次
- Scaling Automated Database System TestingSuyang Zhong, Manuel RiggerASPLOS 2026 · 被引用 4 次
- SymFit: Making the Common (Concrete) Case Fast for Binary-Code Concolic ExecutionZhenxiao Qi, Jie Hu, Zhaoqi Xiao, Heng YinUSENIX Security 2024 · 被引用 4 次
- Sand: Decoupling Sanitization from Fuzzing for Low OverheadZiqiao Kong, Shaohua Li, Heqing Huang, Zhendong SuICSE 2025 · 被引用 1 次
它引用的顶会 Paper15
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- VUzzer: Application-aware Evolutionary FuzzingSanjay Rawat, Vivek Jain, Ashish Kumar, Lucian Cojocar 等NDSS 2017 · 被引用 700 次
- Angora: Efficient Fuzzing by Principled SearchPeng Chen, Hao ChenS&P 2018 · 被引用 616 次
- CollAFL: Path Sensitive FuzzingShuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu 等S&P 2018 · 被引用 426 次
相关 Paper
- FuzzerAid: Grouping Fuzzed Crashes Based On Fault SignaturesAshwin Kallingal Joshy, Wei LeASE 2022 · 被引用 6 次
- Sleuth: A Switchable Dual-Mode Fuzzer to Investigate Bug Impacts Following a Single PoCHaolai Wei, Liwei Chen, Zhijie Zhang, Gang Shi 等ISSTA 2024 · 被引用 1 次
- GPTrace: Effective Crash Deduplication Using LLM EmbeddingsPatrick Herter, Vincent Ahlrichs, Ridvan Açilan, Julian HorschICSE 2026
- An In-depth Analysis of Duplicated Linux Kernel Bug ReportsDongliang Mu, Yuhang Wu, Yueqi Chen, Zhenpeng Lin 等NDSS 2022
- On the Reliability of Coverage-Based Fuzzer BenchmarkingMarcel Böhme, László Szekeres, Jonathan MetzmanICSE 2022 · 被引用 91 次
