SinkFinder: harvesting hundreds of unknown interesting function pairs with just one seed
Pan Bian, Bin Liang, Jianjun Huang, Wenchang Shi, Xidong Wang, Jian Zhang
Abstract
Mastering the knowledge about security-sensitive functions that can potentially result in bugs is valuable to detect them. However, identifying this kind of functions is not a trivial task. Introducing machine learning-based techniques to do the task is a natural choice. Unfortunately, the approach also requires considerable prior knowledge, e.g., sufficient labelled training samples. In practice, the requirement is often hard to meet.
In this paper, to solve the problem, we propose a novel and practical method called SinkFinder to automatically discover function pairs that we are interested in, which only requires very limited prior knowledge. SinkFinder first takes just one pair of wellknown interesting functions as the initial seed to infer enough positive and negative training samples by means of sub-word word embedding. By using these samples, a support vector machine classifier is trained to identify more interesting function pairs. Finally, checkers equipped with the obtained knowledge can be easily developed to detect bugs in target systems. The experiments demonstrate that SinkFinder can successfully discover hundreds of interesting functions and detect dozens of previously unknown bugs from large-scale systems, such as Linux, OpenSSL and PostgreSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fd588de-60ab-4962-8e46-2075e9ecc6d4Cited by top-tier papers6
- Goshawk: Hunting Memory Corruptions via Structure-Aware and Object-Centric Memory Operation SynopsisYunlong Lyu, Yi Fang, Yiwei Zhang, Qibin Sun et al.S&P 2022 · 26 citations
- Detecting Kernel Memory Bugs through Inconsistent Memory Management Intention InferencesDinghao Liu, Zhipeng Lu, Shouling Ji, Kangjie Lu et al.USENIX Security 2024 · 6 citations
- Boosting Static Resource Leak Detection via LLM-based Resource-Oriented Intention InferenceChong Wang, Jianan Liu, Xin Peng, Yang Liu et al.ICSE 2025 · 5 citations
- Detecting Memory Errors in Python Native Code by Tracking Object Lifecycle with Reference CountXutong Ma, Jiwei Yan, Hao Zhang, Jun Yan et al.ASE 2023 · 2 citations
- Uncovering the iceberg from the tip: Generating API Specifications for Bug Detection via Specification Propagation AnalysisMiaoqian Lin, Kai Chen, Yi Yang, Jinghua LiuNDSS 2025
Builds on5
- DIFUZE: Interface Aware Fuzzing for Kernel DriversJake Corina, Aravind Machiry, Christopher Salls, Yan Shoshitaishvili et al.CCS 2017 · 195 citations
- SemFuzz: Semantics-based Automatic Generation of Proof-of-Concept ExploitsWei You, Peiyuan Zong, Kai Chen, XiaoFeng Wang et al.CCS 2017 · 148 citations
- APISan: Sanitizing API Usages through Semantic Cross-CheckingInsu Yun, Changwoo Min, Xujie Si, Yeongjin Jang et al.USENIX Security 2016 · 107 citations
- Automatically Detecting Error Handling Bugs Using Error SpecificationsSuman Jana, Yuan Jochen Kang, Samuel Roth, Baishakhi RayUSENIX Security 2016 · 79 citations
- K-Miner: Uncovering Memory Corruption in LinuxDavid Gens, Simon Schmitt, Lucas Davi, Ahmad-Reza SadeghiNDSS 2018 · 58 citations
Related papers
- Raisin: Identifying Rare Sensitive Functions for Bug DetectionJianjun Huang, Jianglei Nie, Yuanjun Gong, Wei You et al.ICSE 2024 · 2 citations
- Top Score on the Wrong Exam: On Benchmarking in Machine Learning for Vulnerability DetectionNiklas Risse, Jing Liu, Marcel BöhmeISSTA 2025 · 8 citations
- DR. CHECKER: A Soundy Analysis for Linux Kernel DriversAravind Machiry, Chad Spensky, Jake Corina, Nick Stephens et al.USENIX Security 2017 · 126 citations
- Balancing Analysis Time and Bug Detection: Daily Development-friendly Bug Detection in LinuxKeita Suzuki, Kenta Ishiguro, Kenji KonoUSENIX ATC 2024 · 6 citations
- Distinguishing Look-Alike Innocent and Vulnerable Code by Subtle Semantic Representation Learning and ExplanationChao Ni, Xin Yin, Kaiwen Yang, Dehai Zhao et al.FSE 2023 · 42 citations
