Beyond Text-to-SQL: Can LLMs Really Debug Enterprise ETL SQL?
Jing Ye, Yiwen Duan, Yonghong Yu, Victor Ma, Yang Gao, Xing Chen
摘要
SQL is central to enterprise data engineering, yet generating fully correct SQL code in a single attempt remains difficult-even for experienced developers and advanced Text-to-SQL LLMs-often requiring multiple debugging iterations. We introduce Squirrel Benchmark, the first benchmark for enterprise-level SQL reasoning and debugging. Our benchmark is built upon two key innovations: (1) an automated construction workflow that employs reverse engineering to systematically inject realistic bugs into largescale SQL code, enabling scalable and diverse benchmark generation; and (2) an execution-free evaluation framework tailored for enterprise settings, providing fast, accurate, and resourceefficient assessment. Squirrel Benchmark comprises 469 Squirrel-Syntax queries featuring syntax errors with explicit error messages, and 516 Squirrel-Semantic queries targeting semantic errors where codes fails to meet user intent. The queries are highly complex, averaging over 140 lines, and featuring deep and wide abstract syntax trees (average width > 11, depth > 8.7). Evaluation of nearly 30 LLMs reveals a substantial performance gap: the best-performing model, Claude-4-Sonnet, achieves only 36.46% accuracy on Squirrel-Syntax and 32.17% on Squirrel-Semantic, while most models score below 20%. We further explore four solution strategies, identify key challenges, and outline promising directions for enterprise SQL debugging with LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 被引用 491 次
相关 Paper
- Evaluating Cross-Domain Text-to-SQL Models and BenchmarksMohammadreza Pourreza, Davood RafieiEMNLP 2023 · 被引用 14 次
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao 等ICLR 2025
- SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL BenchmarksMohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Oroojlooy, Graham Horwood 等ACL 2026
- NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL SolutionsShizheng Hou, Wenqi Pei, Nuo Chen, Quang-Trung Ta 等VLDB 2026 · 被引用 1 次
- GBV-SQL: Guided Generation and SQL2Text Back-Translation Validation for Multi-Agent Text2SQLDaojun Chen, Xi Wang, Shenyuan Ren, Qingzhi Ma 等ACL 2026
