DuSQL: A Large-Scale and Pragmatic Chinese Text-to-SQL Dataset
Lijie Wang, Ao Zhang, Kun Wu, Ke Sun, Zhenghua Li, Hua Wu, Min Zhang, Haifeng Wang
摘要
Due to the lack of labeled data, previous research on text-to-SQL parsing mainly focuses on English. Representative English datasets include ATIS, WikiSQL, Spider, etc. This paper presents DuSQL, a larges-scale and pragmatic Chinese dataset for the cross-domain text-to-SQL task, containing 200 databases, 813 tables, and 23,797 question/SQL pairs. Our new dataset has three major characteristics. First, by manually analyzing questions from several representative applications, we try to figure out the true distribution of SQL queries in real-life needs. Second, DuSQL contains a considerable proportion of SQL queries involving row or column calculations, motivated by our analysis on the SQL query distributions. Finally, we adopt an effective data construction framework via human-computer collaboration. The basic idea is automatically generating SQL queries based on the SQL grammar and constrained by the given database. This paper describes in detail the construction process and data statistics of DuSQL. Moreover, we present and compare performance of several open-source textto-SQL parsers with minor modification to accommodate Chinese, including a simple yet effective extension to IRNet for handling calculation SQL queries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesBolei Ma, Yuting Li, Wei Zhou, Ziwei Gong 等ACL 2025 · 被引用 28 次
- Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingKun Wu, Lijie Wang, Zhenghua Li, Ao Zhang 等EMNLP 2021 · 被引用 22 次
- Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic KnowledgeLongxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan 等EMNLP 2022 · 被引用 11 次
- RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi 等ACL 2023 · 被引用 7 次
- LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Complex ReasoningLiutao, Xutao Mao, Dixuan Zhang, Yifan Li 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- Chase: A Large-Scale and Pragmatic Chinese Dataset for Cross-Database Context-Dependent Text-to-SQLJiaqi Guo, Ziliang Si, Yu Wang, Qian Liu 等ACL 2021
- MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic ParsingLongxu Dou, Yan Gao, Mingyang Pan, Dingzirui Wang 等AAAI 2023 · 被引用 35 次
- Bridging the Generalization Gap in Text-to-SQL Parsing with Schema ExpansionChen Zhao, Yu Su, Adam Pauls, Emmanouil Antonios PlataniosACL 2022 · 被引用 19 次
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
- SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQLRuichu Cai, Jinjie Yuan, Boyan Xu, Zhifeng HaoNeurIPS 2021 · 被引用 90 次
