Automated Discovery of Test Oracles for Database Management Systems Using LLMs
Qiuyang Mang, Runyuan He, Suyang Zhong, Xiaoxuan Liu, Huanchen Zhang, Alvin Cheung
摘要
Since 2020, automated testing for Database Management Systems (DBMSs) has flourished, uncovering hundreds of bugs in widely-used systems. A cornerstone of these techniques is test oracle, which typically implements a mechanism to generate equivalent query pairs, and subsequently runs the pair and identifies bugs by checking the consistency of their results. While running these oracles can be automated, designing the mechanism to generate equivalent queries remains a fundamentally manual endeavor. This paper explores the use of large language models (LLMs) to automate the discovery of equivalent queries in the design of test oracles, addressing a long-standing bottleneck towards fully automated DBMS testing.
Although LLMs demonstrate impressive creativity, they are prone to hallucinations that can produce numerous false positive bug reports. Furthermore, their high monetary cost and latency mean that LLM invocations should be limited to ensure that bug detection is efficient and economical. To this end, we introduce Argus, a novel framework built upon the core concept of the Constrained Abstract Query-a SQL skeleton containing placeholders and their associated instantiation conditions, e.g., the placeholder must be filled by a Boolean column. Argus uses LLMs to generate pairs of these skeletons, with their equivalence formally proven using a SQL equivalence solver to ensure soundness. After that, the placeholders in the verified skeletons are instantiated with concrete, reusable SQL snippets that are also synthesized by LLMs to produce complex test cases. We have implemented Argus and evaluated it on five extensively tested DBMSs, discovering 41 previously unknown bugs, 36 of which are logic bugs, with 36 confirmed and 27 already fixed by the developers. The artifacts for Argus are available at https://github.com/joyemang33/Argus CCS Concepts: • Information systems → Database management system engines; • Software and its engineering → Software testing and debugging; • Computing methodologies → Machine learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper51
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel 等ICSE 2024 · 被引用 155 次
- Testing Database Engines via Pivoted Query SynthesisManuel Rigger, Zhendong SuOSDI 2020 · 被引用 150 次
- Enhancing Static Analysis for Practical Bug Detection: An LLM-Integrated ApproachHaonan Li, Yu Hao, Yizhuo Zhai, Zhiyun QianOOPSLA 2024 · 被引用 142 次
- Finding bugs in database systems via query partitioningManuel Rigger, Zhendong SuOOPSLA 2020 · 被引用 116 次
相关 Paper
- Scaling Automated Database System TestingSuyang Zhong, Manuel RiggerASPLOS 2026 · 被引用 4 次
- LLMSQLMUTATOR: LLM-Powered Test Case Generation for Database Using Bug ReportsChenglin Tian, Chaofan Li, Yawen Li, Yingxia ShaoICDE 2026
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- SmartFuzz: Leveraging Large Language Models and Feature Composition to Generate High-Quality Seeds for Database FuzzingLi Lin, Jintai Hong, Yanlin Zhuang, Rongxin WuOOPSLA 2026
- ACME: Automated Clause Mapping Engine for Testing Emerging Database SystemsYuancheng Jiang, Jianing Wang, Chuqi Zhang, Roland H. C. Yap 等FSE 2026
