Faithful Low-Resource Data-to-Text Generation through Cycle Training
Zhuoer Wang, Marcus D. Collins, Nikhita Vedula, Simone Filice, Shervin Malmasi, Oleg Rokhlenko
Abstract
Methods to generate text from structured data have advanced significantly in recent years, primarily due to fine-tuning of pre-trained language models on large datasets. However, such models can fail to produce output faithful to the input data, particularly on out-of-domain data. Sufficient annotated data is often not available for specific domains, leading us to seek an unsupervised approach to improve the faithfulness of output text. Since the problem is fundamentally one of consistency between the representations of the structured data and text, we evaluate the effectiveness of cycle training in this work. Cycle training uses two models which are inverses of each other: one that generates text from structured data, and one which generates the structured data from natural language text. We show that cycle training, when initialized with a small amount of supervised data (100 samples in our case), achieves nearly the same performance as fully supervised approaches for the data-to-text generation task on the WebNLG, E2E, WTQ, and WSQL datasets. We perform extensive empirical analysis with automated evaluation metrics and a newly designed human evaluation schema to reveal different cycle training strategies' effectiveness of reducing various types of generation errors. Our code is publicly available at https:// github.com/Edillower/CycleNLG .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b16095c-7fdc-403a-bc59-a82740a9cbc9Builds on8
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
- ToTTo: A Controlled Table-To-Text Generation DatasetAnkur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui et al.EMNLP 2020 · 69 citations
Related papers
- CycleNER: An Unsupervised Training Approach for Named Entity RecognitionAndrea Iovine, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko et al.WWW 2022 · 19 citations
- CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle ConsistencyZhanming Shen, Hao Chen, Yulei Tang, Shaolin Zhu et al.EMNLP 2025
- PLOG: Table-to-Logic Pretraining for Logical Table-to-Text GenerationAo Liu, Haoyu Dong, Naoaki Okazaki, Shi Han et al.EMNLP 2022 · 15 citations
- Unsupervised Domain Adaptation for Referring Semantic SegmentationHaonan Shi, Wenwen Pan, Zhou Zhao, Mingmin Zhang et al.ACM MM 2023 · 5 citations
- Leveraging Unpaired Data for Vision-Language Generative Models via Cycle ConsistencyTianhong Li, Sangnie Bhardwaj, Yonglong Tian, Han Zhang et al.ICLR 2024 · 8 citations
