Search and Learn: Improving Semantic Coverage for Data-to-Text Generation
Shailza Jolly, Zi Xuan Zhang, Andreas Dengel, Lili Mou
摘要
Data-to-text generation systems aim to generate text descriptions based on input data (often represented in the tabular form). A typical system uses huge training samples for learning the correspondence between tables and texts. However, large training sets are expensive to obtain, limiting the applicability of these approaches in real-world scenarios. In this work, we focus on few-shot data-to-text generation. We observe that, while fine-tuned pretrained language models may generate plausible sentences, they suffer from the low semantic coverage problem in the few-shot setting. In other words, important input slots tend to be missing in the generated text. To this end, we propose a search-and-learning approach that leverages pretrained language models but inserts the missing slots to improve the semantic coverage. We further finetune our system based on the search results to smooth out the search noise, yielding better-quality text and improving inference efficiency to a large extent. Experiments show that our model achieves high performance on E2E and WikiBio datasets. Especially, we cover 98.35% of input slots on E2E, largely alleviating the low coverage problem. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Learning Non-Autoregressive Models from Search for Unsupervised Sentence SummarizationPuyuan Liu, Chenyang Huang, Lili MouACL 2022 · 被引用 20 次
- EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine TranslationYuqiao Wen, Behzad Shayegh, Chenyang Huang, Yanshuai Cao 等AAAI 2025 · 被引用 8 次
- CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High QualityLiang Li, Ruiying Geng, Chengyang Fang, Bing Li 等ACL 2023 · 被引用 2 次
它引用的顶会 Paper4
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen 等ACL 2020 · 被引用 116 次
- Unsupervised Paraphrasing by Simulated AnnealingXianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 等ACL 2020 · 被引用 74 次
- Unsupervised Text Generation by Learning from SearchJingjing Li, Zichao Li, Lili Mou, Xin Jiang 等NeurIPS 2020 · 被引用 60 次
- Discrete Optimization for Unsupervised Sentence Summarization with Word-Level ExtractionRaphael Schumann, Lili Mou, Yao Lu, Olga Vechtomova 等ACL 2020 · 被引用 4 次
相关 Paper
- Neural Pipeline for Zero-Shot Data-to-Text GenerationZdenek Kasner, Ondrej DusekACL 2022
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationWenhu Chen, Yu Su, Xifeng Yan, William Yang WangEMNLP 2020 · 被引用 115 次
- The Benefits of Label-Description Training for Zero-Shot Text ClassificationLingyu Gao, Debanjan Ghosh, Kevin GimpelEMNLP 2023 · 被引用 6 次
- Ontology-enhanced Prompt-tuning for Few-shot LearningHongbin Ye, Ningyu Zhang, Shumin Deng, Xiang Chen 等WWW 2022 · 被引用 78 次
- Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance GenerationJiaxin Huang, Yu Meng, Jiawei HanKDD 2022 · 被引用 17 次
