R2D2: Robust Data-to-Text with Replacement Detection
Linyong Nan, Lorenzo Jaime Yu Flores, Yilun Zhao, Yixin Liu, Luke Benson, Weijin Zou, Dragomir Radev
Abstract
Unfaithful text generation is a common problem for text generation systems. In the case of Data-to-Text (D2T) systems, the factuality of the generated text is particularly crucial for any real-world applications. We introduce R2D2, a training framework that addresses unfaithful Data-to-Text generation by training a system both as a generator and a faithfulness discriminator with additional replacement detection and unlikelihood learning tasks. To facilitate such training, we propose two methods for sampling unfaithful sentences. We argue that the poor entity retrieval capability of D2T systems is one of the primary sources of unfaithfulness, so in addition to the existing metrics, we further propose named entity based metrics to evaluate the fidelity of D2T generations. Our experimental results show that R2D2 systems could effectively mitigate the unfaithful text generation, and they achieve new state-of-theart results on FeTaQA, LogicNLG, and ToTTo, all with significant improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Generation of Training Examples for Tabular Natural Language InferenceJean-Flavien Bussotti, Enzo Veltri, Donatello Santoro, Paolo PapottiSIGMOD 2024 · 7 citations
- TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based ReasoningYilun Zhao, Lyuhao Chen, Arman Cohan, Chen ZhaoACL 2024 · 5 citations
- QTSumm: Query-Focused Summarization over Tabular DataYilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi et al.EMNLP 2023 · 4 citations
Builds on20
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
Related papers
- Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric ViewTianyu Liu, Xin Zheng, Baobao Chang, Zhifang SuiAAAI 2021 · 37 citations
- Contrastive Error Attribution for Finetuned Language ModelsFaisal Ladhak, Esin Durmus, Tatsunori HashimotoACL 2023
- Logical Natural Language Generation from Open-Domain TablesWenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen et al.ACL 2020 · 116 citations
- Re-Imagen: Retrieval-Augmented Text-to-Image GeneratorWenhu Chen, Hexiang Hu, Chitwan Saharia, William W. CohenICLR 2023 · 44 citations
- PLOG: Table-to-Logic Pretraining for Logical Table-to-Text GenerationAo Liu, Haoyu Dong, Naoaki Okazaki, Shi Han et al.EMNLP 2022 · 15 citations
