Search and Learn: Improving Semantic Coverage for Data-to-Text Generation
Shailza Jolly, Zi Xuan Zhang, Andreas Dengel, Lili Mou
Abstract
Data-to-text generation systems aim to generate text descriptions based on input data (often represented in the tabular form). A typical system uses huge training samples for learning the correspondence between tables and texts. However, large training sets are expensive to obtain, limiting the applicability of these approaches in real-world scenarios. In this work, we focus on few-shot data-to-text generation. We observe that, while fine-tuned pretrained language models may generate plausible sentences, they suffer from the low semantic coverage problem in the few-shot setting. In other words, important input slots tend to be missing in the generated text. To this end, we propose a search-and-learning approach that leverages pretrained language models but inserts the missing slots to improve the semantic coverage. We further finetune our system based on the search results to smooth out the search noise, yielding better-quality text and improving inference efficiency to a large extent. Experiments show that our model achieves high performance on E2E and WikiBio datasets. Especially, we cover 98.35% of input slots on E2E, largely alleviating the low coverage problem. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Learning Non-Autoregressive Models from Search for Unsupervised Sentence SummarizationPuyuan Liu, Chenyang Huang, Lili MouACL 2022 · 20 citations
- EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine TranslationYuqiao Wen, Behzad Shayegh, Chenyang Huang, Yanshuai Cao et al.AAAI 2025 · 8 citations
- CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High QualityLiang Li, Ruiying Geng, Chengyang Fang, Bing Li et al.ACL 2023 · 2 citations
Builds on4
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou et al.ACL 2020 · 74 citations
- Unsupervised Text Generation by Learning from SearchJingjing Li, Zichao Li, Lili Mou, Xin Jiang et al.NeurIPS 2020 · 60 citations
- Discrete Optimization for Unsupervised Sentence Summarization with Word-Level ExtractionRaphael Schumann, Lili Mou, Yao Lu, Olga Vechtomova et al.ACL 2020 · 4 citations
Related papers
- Neural Pipeline for Zero-Shot Data-to-Text GenerationZdenek Kasner, Ondrej DusekACL 2022
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationWenhu Chen, Yu Su, Xifeng Yan, William Yang WangEMNLP 2020 · 115 citations
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLingyu Gao, Debanjan Ghosh, Kevin GimpelEMNLP 2023 · 6 citations
- Ontology-enhanced Prompt-tuning for Few-shot LearningHongbin Ye, Ningyu Zhang, Shumin Deng, Xiang Chen et al.WWW 2022 · 78 citations
- Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance GenerationJiaxin Huang, Yu Meng, Jiawei HanKDD 2022 · 17 citations
