TGEA: An Error-Annotated Dataset and Benchmark Tasks for TextGeneration from Pretrained Language Models
Jie He, Bo Peng, Yi Liao, Qun Liu, Deyi Xiong
摘要
In order to deeply understand the capability of pretrained language models in text generation and conduct a diagnostic evaluation, we propose TGEA 1 , an error-annotated dataset with multiple benchmark tasks for text generation from pretrained language models (PLMs). We use carefully selected prompt words to guide GPT-2 to generate candidate sentences, from which we select 47K for error annotation. Crowdsourced workers manually check each of these sentences and detect 12k erroneous sentences. We create an error taxonomy to cover 24 types of errors occurring in these erroneous sentences according to the nature of errors with respect to linguistics and knowledge (e.g., common sense). For each erroneous span in PLM-generated sentences, we also detect another span that is closely associated with it. Each error is hence manually labeled with comprehensive annotations, including the span of the error, the associated span, minimal correction to the error, the type of the error, and rationale behind the error. Apart from the fully annotated dataset, we also present a detailed description of the data collection procedure, statistics and analysis of the dataset. This is the first dataset with comprehensive annotations for PLM-generated texts, which facilitates the diagnostic evaluation of PLM-based text generation. Furthermore, we use TGEA as a benchmark dataset and propose a series of automatic diagnosis tasks, including error detection, error type classification, associated span detection, error rationale generation, to further promote future study on the automatic error detection and correction on texts generated by pretrained language models. * Equal Contributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated TextLiam Dugan, Daphne Ippolito, Arun Kirubarajan, Sherry Shi 等AAAI 2023 · 被引用 112 次
- RuCoLA: Russian Corpus of Linguistic AcceptabilityVladislav Mikhailov, Tatiana Shamardina, Max Ryabinin, Alena Pestova 等EMNLP 2022 · 被引用 19 次
它引用的顶会 Paper16
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Retrospective Reader for Machine Reading ComprehensionZhuosheng Zhang, Junjie Yang, Hai ZhaoAAAI 2021 · 被引用 237 次
相关 Paper
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2021
- Targeted Syntactic Evaluation for Grammatical Error CorrectionAomi Koyama, Masato Mita, Su-Youn Yoon, Yasufumi Takama 等ACL 2025
- Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine TextYao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith 等ACL 2022
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsPei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin 等EMNLP 2021 · 被引用 28 次
- A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao 等ACL 2022 · 被引用 194 次
