PLForge: Enhancing Language Models for Natural Language to Procedural Extensions of SQL
Hang Zhang, Chaokun Wang, Hongwei Li, Cheng Wu, Songyao Wang, Yabin Liu, Gengyuan Shi, Ziyang Liu
Abstract
Procedural Language extensions of SQL (abbr. PL/SQL) enhance database programming by integrating procedural constructs with SQL's declarative syntax, thereby improving the reusability, modularity, and maintainability of SQL. Besides, PL/SQL in database systems presents significant challenges in real-world development, primarily due to the inherent complexity of programming. To reduce the development difficulty of PL/SQL, this paper studies the novel task of translating natural language (NL) to PL/SQL (i.e., NL-to-PL/SQL), aimed at simplifying PL/SQL development. Recent advancements in language models have shown promise in translating natural language questions into SQL queries (i.e., Text-to-SQL). However, the state-of-the-art Text-to-SQL methods focus only on single SQL queries, neglecting the procedural extensions of SQL, which limits their effectiveness for the NL-to-PL/SQL task. In this paper, we propose PLForge, a suite of pre-trained language models with parameter configurations of 3B, 7B, and 15B, tailored for NL-to-PL/SQL tasks. To enhance the PL/SQL generation capabilities of PLForge, we leverage a curated PL/SQL-centric data corpus and employ an incremental pre-training approach. Furthermore, to fully exploit the potential of PLForge, we propose a comprehensive prompt construction strategy tailored specifically for PL/SQL. Given the scarcity of NL-to-PL/SQL datasets, we develop a template-based method for generating NL-to-PL/SQL data. We conduct a series of experiments on PLForge and several baseline models. Based on execution match and exact match metrics that are designed specifically for the NL-to-PL/SQL task, the experimental results demonstrate that PLForge outperforms existing models in both in-context learning and supervised fine-tuning settings.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 35a8ea0b-ba25-47ed-ba0b-0c3423722bd7Cited by top-tier papers3
- Negative Feedback Really Matters: Signed Dual-Channel Graph Contrastive Learning Framework for RecommendationLeqi Zheng, Chaokun Wang, Zixin Song, Cheng Wu et al.NeurIPS 2025 · 6 citations
- What Should I Cite? A RAG Benchmark for Academic Citation PredictionLeqi Zheng, Jiajun Zhang, Canzhi Chen, Chaokun Wang et al.WWW 2026 · 2 citations
- Adaptive Text2GQL: Integrating Structural Twig Linking and Evolutionary In-Context LearningFang Niu, Chaokun Wang, Hang Zhang, Songyao WangACL 2026
Related papers
- Structure-Guided Large Language Models for Text-to-SQL GenerationQinggang Zhang, Hao Chen, Junnan Dong, Shengyuan Chen et al.ICML 2025
- MIGA: A Unified Multi-Task Generation Framework for Conversational Text-to-SQLYingwen Fu, Wenjie Ou, Zhou Yu, Yue LinAAAI 2023 · 14 citations
- PURPLE: Making a Large Language Model a Better SQL WriterTonghui Ren, Yuankai Fan, Zhenying He, Ren Huang et al.ICDE 2024 · 49 citations
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 9 citations
