KGPT: Knowledge-Grounded Pre-Training for Data-to-Text Generation
Wenhu Chen, Yu Su, Xifeng Yan, William Yang Wang
Abstract
Data-to-text generation has recently attracted substantial interests due to its wide applications. Existing methods have shown impressive performance on an array of tasks. However, they rely on a significant amount of labeled data for each task, which is costly to acquire and thus limits their application to new tasks and domains. In this paper, we propose to leverage pre-training and transfer learning to address this issue. We propose a knowledge-grounded pre-training (KGPT), which consists of two parts, 1) a general knowledge-grounded generation model to generate knowledge-enriched text. 2) a pre-training paradigm on a massive knowledge-grounded text corpus crawled from the web. The pre-trained model can be fine-tuned on various data-to-text generation tasks to generate task-specific text. We adopt three settings, namely fully-supervised, zero-shot, few-shot to evaluate its effectiveness. Under the fully-supervised setting, our model can achieve remarkable gains over the known baselines. Under zero-shot setting, our model without seeing any examples achieves over 30 ROUGE-L on WebNLG while all other baselines fail. Under the few-shot setting, our model only needs about one-fifteenth as many labeled examples to achieve the same level of performance as baseline models. These experiments consistently prove the strong generalization ability of our proposed framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ae25b71-2713-416a-b88e-ba314c8a6c67Cited by top-tier papers23
- KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsJifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao et al.ICLR 2024 · 91 citations
- Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language ModelsShangbin Feng, Weijia Shi, Yuyang Bai, Vidhisha Balachandran et al.ICLR 2024 · 56 citations
- Open Domain Question Answering with A Unified Knowledge InterfaceKaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg et al.ACL 2022 · 45 citations
- Knowledge-Grounded Self-Rationalization via Extractive and Natural Language ExplanationsBodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, Julian J. McAuleyICML 2022 · 40 citations
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung et al.SIGMOD 2026 · 17 citations
Builds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
- Variational Template Machine for Data-to-Text GenerationRong Ye, Wenxian Shi, Hao Zhou, Zhongyu Wei et al.ICLR 2020 · 45 citations
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno et al.ACL 2020 · 19 citations
Related papers
- KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationQi Zhu, Fei Mi, Zheng Zhang, Yasheng Wang et al.AAAI 2023 · 5 citations
- Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text ClassificationShengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu et al.ACL 2022
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang et al.AAAI 2024 · 7 citations
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang et al.KDD 2022 · 141 citations
- Edge Prompt Tuning for Graph Neural NetworksXingbo Fu, Yinhan He, Jundong LiICLR 2025 · 140 citations
