Non-Programmers Can Label Programs Indirectly via Active Examples: A Case Study with Text-to-SQL
Ruiqi Zhong, Charlie Snell, Dan Klein, Jason Eisner
Abstract
Can non-programmers annotate natural language utterances with complex programs that represent their meaning? We introduce APEL, a framework in which non-programmers select among candidate programs generated by a seed semantic parser (e.g., Codex). Since they cannot understand the candidate programs, we ask them to select indirectly by examining the programs’ input-ouput examples. For each utterance, APEL actively searches for a simple input on which the candidate programs tend to produce different outputs. It then asks the non-programmers only to choose the appropriate output, thus allowing us to infer which program is correct and could be used to fine-tune the parser. As a first case study, we recruited human non-programmers to use APEL to re-annotate SPIDER, a text-to-SQL dataset. Our approach achieved the same annotation accuracy as the original expert annotators (75%) and exposed many subtle errors in the original annotations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsHelia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme et al.ACL 2024 · 27 citations
- SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQLGe Qu, Jinyang Li, Bowen Qin, Xiaolong Li et al.ACL 2025 · 13 citations
- The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and TrustNishant Subramani, Palash Goyal, Yiwen Song, Mani Malek et al.ICML 2026 · 1 citation
- COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto FrontierGaoxiang Luo, Aryan DeshwalEMNLP 2025
- Language Models Learn to Mislead Humans via RLHFJiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez et al.ICLR 2025
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
Related papers
- Automated Repair of Programs from Large Language ModelsZhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury et al.ICSE 2023 · 213 citations
- "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQLYuntao Li, Bei Chen, Qian Liu, Yan Gao et al.EMNLP 2020 · 17 citations
- SPARQLing Database Queries from Intermediate Question DecompositionsIrina Saparina, Anton OsokinEMNLP 2021 · 11 citations
- Speak to your Parser: Interactive Text-to-SQL with Natural Language FeedbackAhmed Elgohary, Saghar Hosseini, Ahmed Hassan AwadallahACL 2020 · 13 citations
- Parsel🦆: Algorithmic Reasoning with Language Models by Composing DecompositionsEric Zelikman, Qian Huang, Gabriel Poesia, Noah D. Goodman et al.NeurIPS 2023 · 90 citations
