Synthetic Programming Elicitation for Text-to-Code in Very Low-Resource Programming and Formal Languages
Federico Mora, Justin Wong, Haley Lepe, Sahil Bhatia, Karim Elmaaroufi, George Varghese, Joseph E. Gonzalez, Elizabeth Polgreen, Sanjit A. Seshia
Abstract
Recent advances in large language models (LLMs) for code applications have demonstrated remarkable zero-shot fluency and instruction following on challenging code related tasks ranging from test case generation to self-repair. Unsurprisingly, however, models struggle to compose syntactically valid programs in programming languages unrepresented in pre-training, referred to as very low-resource Programming Languages (VLPLs). VLPLs appear in crucial settings, including domain-specific languages for internal tools, tool-chains for legacy languages, and formal verification frameworks. Inspired by a technique called natural programming elicitation, we propose designing an intermediate language that LLMs"naturally"know how to use and which can be automatically compiled to a target VLPL. When LLMs generate code that lies outside of this intermediate language, we use compiler techniques to repair the code into programs in the intermediate language. Overall, we introduce synthetic programming elicitation and compilation (SPEAC), an approach that enables LLMs to generate syntactically valid code even for VLPLs. We empirically evaluate the performance of SPEAC in a case study for the UCLID5 formal verification language and find that, compared to existing retrieval and fine-tuning baselines, SPEAC produces syntactically correct programs more frequently and without sacrificing semantic correctness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9c7f41d-69c6-4115-936c-893ba5a1abc5Cited by top-tier papers4
- Verified Code Transpilation with LLMsSahil Bhatia, Jie Qiu, Niranjan Hasabnis, Sanjit A. Seshia et al.NeurIPS 2024 · 36 citations
- QLCoder: A Query Synthesizer For Static Analysis of Security VulnerabilitiesClaire Wang, Ziyang Li, Saikat Dutta, Mayur NaikICLR 2026 · 22 citations
- TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace PartitioningSheng Wang, Pengan Chen, Jingqi Zhou, Qintong Li et al.NeurIPS 2025 · 9 citations
- Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case StudyMingwei Liu, Zheng Pei, Yanlin Wang, Zihao Wang et al.FSE 2026
Builds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Grounded Copilot: How Programmers Interact with Code-Generating ModelsShraddha Barke, Michael B. James, Nadia PolikarpovaOOPSLA 2023 · 408 citations
Related papers
- Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse PromptsXuan-Phi Nguyen, Mahani Aljunied, Shafiq Joty, Lidong BingACL 2024
- Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMsFederico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger et al.OOPSLA 2024 · 33 citations
- HoarePrompt: Structural Reasoning About Program Correctness in Natural LanguageDimitrios Stamatios Bouras, Yihan Dai, Tairan Wang, Yingfei Xiong et al.ICSE 2026
- Examining Zero-Shot Vulnerability Repair with Large Language ModelsHammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri et al.S&P 2023
- ExploraCoder: Advancing Code Generation for Multiple Unseen APIs via Planning and Chained ExplorationYunkun Wang, Yue Zhang, Zhen Qin, Chen Zhi et al.ACL 2025 · 13 citations
