PLForge: Enhancing Language Models for Natural Language to Procedural Extensions of SQL
Hang Zhang, Chaokun Wang, Hongwei Li, Cheng Wu, Songyao Wang, Yabin Liu, Gengyuan Shi, Ziyang Liu
摘要
Procedural Language extensions of SQL (abbr. PL/SQL) enhance database programming by integrating procedural constructs with SQL's declarative syntax, thereby improving the reusability, modularity, and maintainability of SQL. Besides, PL/SQL in database systems presents significant challenges in real-world development, primarily due to the inherent complexity of programming. To reduce the development difficulty of PL/SQL, this paper studies the novel task of translating natural language (NL) to PL/SQL (i.e., NL-to-PL/SQL), aimed at simplifying PL/SQL development. Recent advancements in language models have shown promise in translating natural language questions into SQL queries (i.e., Text-to-SQL). However, the state-of-the-art Text-to-SQL methods focus only on single SQL queries, neglecting the procedural extensions of SQL, which limits their effectiveness for the NL-to-PL/SQL task. In this paper, we propose PLForge, a suite of pre-trained language models with parameter configurations of 3B, 7B, and 15B, tailored for NL-to-PL/SQL tasks. To enhance the PL/SQL generation capabilities of PLForge, we leverage a curated PL/SQL-centric data corpus and employ an incremental pre-training approach. Furthermore, to fully exploit the potential of PLForge, we propose a comprehensive prompt construction strategy tailored specifically for PL/SQL. Given the scarcity of NL-to-PL/SQL datasets, we develop a template-based method for generating NL-to-PL/SQL data. We conduct a series of experiments on PLForge and several baseline models. Based on execution match and exact match metrics that are designed specifically for the NL-to-PL/SQL task, the experimental results demonstrate that PLForge outperforms existing models in both in-context learning and supervised fine-tuning settings.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Negative Feedback Really Matters: Signed Dual-Channel Graph Contrastive Learning Framework for RecommendationLeqi Zheng, Chaokun Wang, Zixin Song, Cheng Wu 等NeurIPS 2025 · 被引用 6 次
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang 等WWW 2026 · 被引用 2 次
- Adaptive Text2GQL: Integrating Structural Twig Linking and Evolutionary In-Context LearningFang Niu, Chaokun Wang, Hang Zhang, Songyao WangACL 2026
相关 Paper
- Structure-Guided Large Language Models for Text-to-SQL GenerationQinggang Zhang, Hao Chen, Junnan Dong, Shengyuan Chen 等ICML 2025
- MIGA: A Unified Multi-Task Generation Framework for Conversational Text-to-SQLYingwen Fu, Wenjie Ou, Zhou Yu, Yue LinAAAI 2023 · 被引用 14 次
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang 等ICDE 2024 · 被引用 49 次
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
