Variational Template Machine for Data-to-Text Generation
Rong Ye, Wenxian Shi, Hao Zhou, Zhongyu Wei, Lei Li
Abstract
How to generate descriptions from structured data organized in tables? Existing approaches using neural encoder-decoder models often suffer from lacking diversity. We claim that an open set of templates is crucial for enriching the phrase constructions and realizing varied generations. Learning such templates is prohibitive since it often requires a large paired <table,description> corpus, which is seldom available. This paper explores the problem of automatically learning reusable "templates" from paired and non-paired data. We propose the variational template machine (VTM), a novel method to generate text descriptions from data tables. Our contributions include: a) we carefully devise a specific model architecture and losses to explicitly disentangle text template and semantic content information in the latent spaces, and b) we utilize both small parallel data and large raw text without aligned tables to enrich the template learning. Experiments on datasets from a variety of different domains show that VTM is able to generate more diversely while keeping a good fluency and quality. * Work done while Rong Ye was a research intern at ByteDance AI Lab. 1 An infobox is a table containing attribute-value data about a certain subject. It is mostly used on Wikipedia pages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 026b67fd-8115-40f9-82ef-2b0d97f403feCited by top-tier papers7
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao et al.ACL 2022 · 194 citations
- Calliope: Automatic Visual Data Story Generation from a SpreadsheetDanqing Shi, Xinyue Xu, Fuling Sun, Yang Shi et al.IEEE VIS 2020 · 179 citations
- KGPT: Knowledge-Grounded Pre-Training for Data-to-Text GenerationWenhu Chen, Yu Su, Xifeng Yan, William Yang WangEMNLP 2020 · 115 citations
- Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric ViewTianyu Liu, Xin Zheng, Baobao Chang, Zhifang SuiAAAI 2021 · 37 citations
- HTKG: Deep Keyphrase Generation with Neural Hierarchical Topic GuidanceYuxiang Zhang, Tao Jiang, Tianyu Yang, Xiaoli Li et al.SIGIR 2022 · 14 citations
Builds on1
Related papers
- TempTabQA: Temporal Question Answering for Semi-Structured TablesVivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang et al.EMNLP 2023 · 4 citations
- INFOTABS: Inference on Tables as Semi-structured DataVivek Gupta, Maitrey Mehta, Pegah Nokhiz, Vivek SrikumarACL 2020
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
- Towards Table-to-Text Generation with Numerical ReasoningLya Hulliyyatus Suadaa, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura et al.ACL 2021
