Syntactically Rich Discriminative Training: An Effective Method for Open Information Extraction
Frank Mtumbuka, Thomas Lukasiewicz
摘要
Open information extraction (OIE) is the task of extracting facts "(Subject, Relation, Object)" from natural language text. We propose several new methods for training neural OIE models in this paper. First, we propose a novel method for computing syntactically rich text embeddings using the structure of dependency trees. Second, we propose a new discriminative training approach to OIE in which tokens in the generated fact are classified as "real" or "fake", i.e., those tokens that are in both the generated and gold tuples, and those that are only in the generated tuple but not in the gold tuple. We also address the issue of repetitive tokens in generated facts and improve the models' ability to generate implicit facts. Our approach reduces repetitive tokens by a factor of 23%. Finally, we present paraphrased versions of the CaRB, OIE2016, and LSOIE datasets, and show that the models' performance substantially improves when trained on datasets augmented by such data. Our best model beats the SOTA of IMoJIE on the recent CaRB dataset, with an improvement of 39.63% in F 1 score.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Span Model for Open Information Extraction on Accurate CorpusJunlang Zhan, Hai ZhaoAAAI 2020 · 被引用 90 次
- Systematic Comparison of Neural Architectures and Training Approaches for Open Information ExtractionPatrick Hohenecker, Frank Mtumbuka, Vid Kocijan, Thomas LukasiewiczEMNLP 2020 · 被引用 10 次
- IMoJIE: Iterative Memory-Based Joint Open Information ExtractionKeshav Kolluru, Samarth Aggarwal, Vipul Rathore, Mausam 等ACL 2020 · 被引用 5 次
相关 Paper
- DetIE: Multilingual Open Information Extraction Inspired by Object DetectionMichael Vasilkovsky, Anton Alekseev, Valentin Malykh, Ilya Shenbin 等AAAI 2022 · 被引用 24 次
- IELM: An Open Information Extraction Benchmark for Pre-Trained Language ModelsChenguang Wang, Xiao Liu, Dawn SongEMNLP 2022 · 被引用 3 次
- Maximal Clique Based Non-Autoregressive Open Information ExtractionBowen Yu, Yucheng Wang, Tingwen Liu, Hongsong Zhu 等EMNLP 2021 · 被引用 14 次
- Open Information Extraction via ChunksKuicai Dong, Aixin Sun, Jung-Jae Kim, Xiaoli LiEMNLP 2023 · 被引用 4 次
- Syntactic Multi-view Learning for Open Information ExtractionKuicai Dong, Aixin Sun, Jung-Jae Kim, Xiaoli LiEMNLP 2022 · 被引用 6 次
