Neural Pipeline for Zero-Shot Data-to-Text Generation
Zdenek Kasner, Ondrej Dusek
Abstract
In data-to-text (D2T) generation, training on in-domain data leads to overfitting to the data representation and repeating training data noise. We examine how to avoid finetuning pretrained language models (PLMs) on D2T generation datasets while still taking advantage of surface realization capabilities of PLMs. Inspired by pipeline approaches, we propose to generate text by transforming single-item descriptions with a sequence of modules trained on general-domain text-based operations: ordering, aggregation, and paragraph compression. We train PLMs for performing these operations on a synthetic corpus WikiFluent which we build from English Wikipedia. Our experiments on two major triple-to-text datasets—WebNLG and E2E—show that our approach enables D2T generation from RDF triples in zero-shot settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e56ba98-d190-46d7-8199-818ac430ca6bCited by top-tier papers4
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 4 citations
- Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma et al.ACL 2023 · 2 citations
- PixT3: Pixel-based Table-To-Text GenerationIñigo Alonso, Eneko Agirre, Mirella LapataACL 2024 · 2 citations
- Language-Guided Music Recommendation for Video via Prompt AnalogiesDaniel McKee, Justin Salamon, Josef Sivic, Bryan C. RussellCVPR 2023
Builds on14
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationWenhu Chen, Yu Su, Xifeng Yan, William Yang WangEMNLP 2020 · 115 citations
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong et al.ACL 2020 · 103 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
Related papers
- Search and Learn: Improving Semantic Coverage for Data-to-Text GenerationShailza Jolly, Zi Xuan Zhang, Andreas Dengel, Lili MouAAAI 2022 · 15 citations
- AggGen: Ordering and Aggregating while GeneratingXinnuo Xu, Ondrej Dusek, Verena Rieser, Ioannis KonstasACL 2021
- Grounding Language Models to Images for Multimodal Inputs and OutputsJing Yu Koh, Ruslan Salakhutdinov, Daniel FriedICML 2023 · 160 citations
- Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text GenerationZdenek Kasner, Ondrej DusekACL 2024 · 10 citations
- Learning to Prompt with Text Only Supervision for Vision-Language ModelsMuhammad Uzair Khattak, Muhammad Ferjad Naeem, Muzammal Naseer, Luc Van Gool et al.AAAI 2025 · 52 citations
