Sedar: Obtaining High-Quality Seeds for DBMS Fuzzing via Cross-DBMS SQL Transfer
Jingzhou Fu, Jie Liang, Zhiyong Wu, Yu Jiang
Abstract
Effective DBMS fuzzing relies on high-quality initial seeds, which serve as the starting point for mutation. These initial seeds should incorporate various DBMS features to explore the state space thoroughly. While built-in test cases are typically used as initial seeds, many DBMSs lack comprehensive test cases, making it difficult to apply state-of-the-art fuzzing techniques directly. To address this, we propose Sedar which produces initial seeds for a target DBMS by transferring test cases from other DBMSs. The underlying insight is that many DBMSs share similar functionalities, allowing seeds that cover deep execution paths in one DBMS to be adapted for other DBMSs. The challenge lies in converting these seeds to a format supported by the grammar of the target database. Sedar follows a three-step process to generate seeds. First, it executes existing SQL test cases within the DBMS they were designed for and captures the schema information during execution. Second, it utilizes large language models (LLMs) along with the captured schema information to guide the generation of new test cases based on the responses from the LLM. Lastly, to ensure that the test cases can be properly parsed and mutated by fuzzers, Sedar temporarily comments out unparsable sections for the fuzzers and uncomments them after mutation. We integrate Sedar into the DBMS fuzzers Sqirrel and Griffin, targeting DBMSs such as Virtuoso, Mon-etDB, DuckDB, and ClickHouse. Evaluation results demonstrate significant improvements in both fuzzers. Specifically, compared to Sqirrel and Griffin with non-transferred seeds, Sedar enhances code coverage by 72. 46%-214.84% and 21.40%-194.46%; compared to Sqirrel and Griffin with native test cases of these DBMSs as initial seeds, incorporating the transferred seeds of Sedar results in an improvement in code coverage by 4.90%-16.20% and 9.73%-28.41%. Moreover, Sedar discovered 70 new vulnerabilities, with 60 out of them being uniquely found by Sedar with transferred seeds, and 19 of them have been assigned with CVEs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54f92641-aeb9-4b81-bf07-72ca054c316eCited by top-tier papers10
- Detecting Metadata-Related Logic Bugs in Database Systems via Raw Database ConstructionJiansen Song, Wensheng Dou, Yu Gao, Ziyu Cui et al.VLDB 2024 · 13 citations
- Understanding and Detecting SQL Function Bugs: Using Simple Boundary Arguments to Trigger Hundreds of DBMS BugsJingzhou Fu, Jie Liang, Zhiyong Wu, Yanyang Zhao et al.EuroSys 2025 · 6 citations
- Detecting Schema-Related Logic Bugs in Relational DBMSs via Equivalent Database ConstructionJiansen Song, Wensheng Dou, Yingying Zheng, Yu Gao et al.VLDB 2025 · 6 citations
- Understanding and Reusing Test Suites Across Database SystemsSuyang Zhong, Manuel RiggerSIGMOD 2025 · 4 citations
- Towards More Complete Constraints for Deep Learning Library Testing via Complementary Set Guided RefinementGwihwan Go, Chijin Zhou, Quan Zhang, Xiazijian Zou et al.ISSTA 2024 · 2 citations
Builds on19
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 382 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- MoonShine: Optimizing OS Fuzzer Seed Selection with Trace DistillationShankara Pailoor, Andrew Aday, Suman JanaUSENIX Security 2018 · 180 citations
Related papers
- SmartFuzz: Leveraging Large Language Models and Feature Composition to Generate High-Quality Seeds for Database FuzzingLi Lin, Jintai Hong, Yanlin Zhuang, Rongxin WuOOPSLA 2026
- SQUIRREL: Testing Database Management Systems with Language Validity and Coverage FeedbackRui Zhong, Yongheng Chen, Hong Hu, Hangfan Zhang et al.CCS 2020 · 5 citations
- Griffin : Grammar-Free DBMS FuzzingJingzhou Fu, Jie Liang, Zhiyong Wu, Mingzhe Wang et al.ASE 2022 · 44 citations
- VIREO: Human-in-the-Loop DBMS Fuzzing with Visualization and LLM SupportJie Liang, Zhiyong Wu, Jingzhou Fu, Chi Zhang et al.ICDE 2026
- DynSQL: Stateful Fuzzing for Database Management Systems with Complex and Valid SQL Query GenerationZu-Ming Jiang, Jia-Ju Bai, Zhendong SuUSENIX Security 2023
