Cracking SQL Barriers: An LLM-based Dialect Translation System
Wei Zhou, Yuyang Gao, Xuanhe Zhou, Guoliang Li
摘要
Automatic dialect translation reduces the complexity of database migration, which is crucial for applications interacting with multiple database systems. However, rule-based translation tools (e.g., SQLGlot, jOOQ, SQLines) are labor-intensive to develop and often (1) fail to translate certain operations, (2) produce incorrect translations due to rule deficiencies, and (3) generate translations compatible with some database versions but not the others.
In this paper, we investigate the problem of automating dialect translation with large language models (LLMs). There are three main challenges. First, queries often involve lengthy content (e.g., excessive column values) and multiple syntax elements that require translation, increasing the risk of LLM hallucination. Second, database dialects have diverse syntax trees and specifications, making it difficult for cross-dialect syntax matching. Third, dialect translation often involves complex many-to-one relationships between source and target operations, making it impractical to translate each operation in isolation. To address these challenges, we propose an automatic dialect translation system CrackSQL. First, we propose Functionality-based Query Processing that segments the query by functionality syntax trees and simplifies the query via (𝑖) customized function normalization and (𝑖𝑖) translation-irrelevant query abstraction. Second, we design a Cross-Dialect Syntax Embedding Model to generate embeddings by the syntax trees and specifications (of certain version), enabling accurate query syntax matching. Third, we propose a Local-to-Global Dialect Translation strategy, which restricts LLM-based translation and validation on operations that cause local failures, iteratively extending these operations until translation succeeds. Experiments show CrackSQL significantly outperforms existing methods (e.g., by up to 77.42%). The code is available at https:// github.com/ weAIDB/ CrackSQL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DBAIOps: A Reasoning LLM-Enhanced Database Operation and Maintenance System using Knowledge GraphsWei Zhou, Peng Sun, Xuanhe Zhou, Qianglei Zang 等VLDB 2026 · 被引用 9 次
- Automating Database-Native Function Code Synthesis with LLMsWei Zhou, Xuanhe Zhou, Qikang He, Guoliang Li 等SIGMOD 2026 · 被引用 6 次
- Dial: A Knowledge-Grounded Dialect-Specific NL2SQL SystemXiang Zhang, Le Zhou, Hongming Xu, Wei Zhou 等VLDB 2026 · 被引用 2 次
- AGRAG: Advanced Graph-Based Retrieval-Augmented Generation for LLMsYubo Wang, Haoyang Li, Fei Teng, Lei ChenICDE 2026
- RISE: Rule-Driven SQL Dialect Translation via Query ReductionXudong Xie, Yuwei Zhang, Wensheng Dou, Yu Gao 等ICSE 2026
它引用的顶会 Paper8
- A Learned Query Rewrite System using Monte Carlo Tree SearchXuanhe Zhou, Guoliang Li, Chengliang Chai, Jianhua FengVLDB 2022 · 被引用 85 次
- GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian OptimizationJiale Lao, Yibo Wang, Yufei Li, Jianping Wang 等VLDB 2024 · 被引用 76 次
- LLM-R2: A Large Language Model Enhanced Rule-based Rewrite System for Boosting Query EfficiencyZhaodonghui Li, Haitao Yuan, Huiming Wang, Gao Cong 等VLDB 2025 · 被引用 52 次
- D-Bot: Database Diagnosis System using Large Language ModelsXuanhe Zhou, Guoliang Li, Zhaoyan Sun, Zhiyuan Liu 等VLDB 2024 · 被引用 50 次
- WeTune: Automatic Discovery and Verification of Query Rewrite RulesZhaoguo Wang, Zhou Zhou, Yicun Yang, Haoran Ding 等SIGMOD 2022 · 被引用 35 次
相关 Paper
- DLBench: A Comprehensive Benchmark for SQL Translation with Large Language ModelsLi Lin, Hongqiao Chen, Qinglin Zhu, Liehang Chen 等ASE 2025
- Dialect-SQL: An Adaptive Framework for Bridging the Dialect Gap in Text-to-SQLJie Shi, Xi Cao, Bo Xu, Jiaqing Liang 等EMNLP 2025 · 被引用 2 次
- Dialect-Agnostic SQL Parsing via LLM-Based SegmentationJunwen An, Kabilan Mahathevan, Manuel RiggerSIGMOD 2026
- LLMSQLMUTATOR: LLM-Powered Test Case Generation for Database Using Bug ReportsChenglin Tian, Chaofan Li, Yawen Li, Yingxia ShaoICDE 2026
- Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating CodeRangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar 等ICSE 2024 · 被引用 96 次
