Automated Discovery of Test Oracles for Database Management Systems Using LLMs
Qiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu, Huanchen Zhang, Alvin Cheung
Abstract
Since 2020, automated testing for Database Management Systems (DBMSs) has flourished, uncovering hundreds of bugs in widely-used systems. A cornerstone of these techniques is test oracle, which typically implements a mechanism to generate equivalent query pairs, and subsequently runs the pair and identifies bugs by checking the consistency of their results. While running these oracles can be automated, designing the mechanism to generate equivalent queries remains a fundamentally manual endeavor. This paper explores the use of large language models (LLMs) to automate the discovery of equivalent queries in the design of test oracles, addressing a long-standing bottleneck towards fully automated DBMS testing.
Although LLMs demonstrate impressive creativity, they are prone to hallucinations that can produce numerous false positive bug reports. Furthermore, their high monetary cost and latency mean that LLM invocations should be limited to ensure that bug detection is efficient and economical. To this end, we introduce Argus, a novel framework built upon the core concept of the Constrained Abstract Query-a SQL skeleton containing placeholders and their associated instantiation conditions, e.g., the placeholder must be filled by a Boolean column. Argus uses LLMs to generate pairs of these skeletons, with their equivalence formally proven using a SQL equivalence solver to ensure soundness. After that, the placeholders in the verified skeletons are instantiated with concrete, reusable SQL snippets that are also synthesized by LLMs to produce complex test cases. We have implemented Argus and evaluated it on five extensively tested DBMSs, discovering 41 previously unknown bugs, 36 of which are logic bugs, with 36 confirmed and 27 already fixed by the developers. The artifacts for Argus are available at https://github.com/joyemang33/Argus CCS Concepts: • Information systems → Database management system engines; • Software and its engineering → Software testing and debugging; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baa3b926-cd88-4448-9cfa-7434e5a39684Cited by top-tier papers1
Ask how each one uses itBuilds on51
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Testing Database Engines via Pivoted Query SynthesisManuel Rigger, Zhendong SuOSDI 2020 · 150 citations
- Enhancing Static Analysis for Practical Bug Detection: An LLM-Integrated ApproachHaonan Li, Yu Hao, Yizhuo Zhai, Zhiyun QianOOPSLA 2024 · 142 citations
- Finding bugs in database systems via query partitioningManuel Rigger, Zhendong SuOOPSLA 2020 · 116 citations
Related papers
- Scaling Automated Database System TestingSuyang Zhong, Manuel RiggerASPLOS 2026 · 4 citations
- LLMSQLMUTATOR: LLM-Powered Test Case Generation for Database Using Bug ReportsChenglin Tian, Chaofan Li, Yawen Li, Yingxia ShaoICDE 2026
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- SmartFuzz: Leveraging Large Language Models and Feature Composition to Generate High-Quality Seeds for Database FuzzingLi Lin, Jintai Hong, Yanlin Zhuang, Rongxin WuOOPSLA 2026
- ACME: Automated Clause Mapping Engine for Testing Emerging Database SystemsYuancheng Jiang, Jianing Wang, Chuqi Zhang, Roland H. C. Yap et al.FSE 2026
