TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
Xiang Huang, Jiayu Shen, Shanshan Huang, Sitao Cheng, Xiaxia Wang, Yuzhong Qu
Abstract
Semantic parsing, which converts natural language questions into logic forms, plays a crucial role in reasoning within structured environments. However, existing methods encounter two significant challenges: reliance on extensive manually annotated datasets and limited generalization capability to unseen examples. To tackle these issues, we propose Targeted Synthetic Data Generation (TARGA), a practical framework that dynamically generates high-relevance synthetic data without manual annotation. Starting from the pertinent entities and relations of a given question, we probe for the potential relevant queries through layer-wise expansion and cross-layer combination. Then we generate corresponding natural language questions for these constructed queries to jointly serve as the synthetic demonstrations for in-context learning. Experiments on multiple knowledge base question answering (KBQA) datasets demonstrate that TARGA, using only a 7B-parameter model, substantially outperforms existing non-finetuned methods that utilize close-sourced model, achieving notable improvements in F1 scores on GrailQA (+7.7) and KBQA-Agent (+12.2). Furthermore, TARGA also exhibits superior sample efficiency, robustness, and generalization capabilities under non-I.I.D. settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfa567d7-7686-4289-b744-ddd4fd31598dBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler et al.WWW 2021 · 304 citations
Related papers
- Code-Style In-Context Learning for Knowledge-Based Question AnsweringZhijie Nie, Richong Zhang, Zhongyuan Wang, Xudong LiuAAAI 2024 · 24 citations
- Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form GenerationShengxiang Gao, Jey Han Lau, Jianzhong QiEMNLP 2025
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question AnsweringXin Sun, Zhongqi Chen, Xing Zheng, Bowen Song et al.ICML 2026 · 2 citations
- Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional GeneralizationInbar Oren, Jonathan Herzig, Jonathan BerantEMNLP 2021
- Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and DatabaseYe Liu, Semih Yavuz, Rui Meng, Dragomir Radev et al.EMNLP 2022 · 21 citations
