Alignment-Augmented Consistent Translation for Multilingual Open Information Extraction
Keshav Kolluru, Muqeeth Mohammed, Shubham Mittal, Soumen Chakrabarti, Mausam
摘要
Progress with supervised Open Information Extraction (OpenIE) has been primarily limited to English due to the scarcity of training data in other languages. In this paper, we explore techniques to automatically convert English text for training OpenIE systems in other languages. We introduce the Alignment-Augmented Constrained Translation (AACTrans) model to translate English sentences and their corresponding extractions consistently with each other — with no changes to vocabulary or semantic meaning which may result from independent translations. Using the data generated with AACTrans, we train a novel two-stage generative OpenIE model, which we call Gen2OIE, that outputs for each sentence: 1) relations in the first stage and 2) all extractions containing the relation in the second stage. Gen2OIE increases relation coverage using a training data transformation technique that is generalizable to multiple languages, in contrast to existing models that use an English-specific training loss. Evaluations on 5 languages — Spanish, Portuguese, Chinese, Hindi and Telugu — show that the Gen2OIE with AACTrans data outperforms prior systems by a margin of 6-25% in F1.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Multilingual Relation Classification via Efficient and Effective PromptingYuxuan Chen, David Harbecke, Leonhard HennigEMNLP 2022 · 被引用 13 次
- Learning to Extract Structured Entities Using Language ModelsHaolun Wu, Ye Yuan, Liana Mikaelyan, Alexander Meulemans 等EMNLP 2024 · 被引用 5 次
- Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social MediaShubham Mittal, Megha Sundriyal, Preslav NakovEMNLP 2023 · 被引用 4 次
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 被引用 4 次
- "Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun CompoundsKeshav Kolluru, Gabriel Stanovsky, MausamEMNLP 2022 · 被引用 1 次
它引用的顶会 Paper4
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield 等ACL 2020 · 被引用 132 次
- X-SRL: A Parallel Cross-Lingual Semantic Role Labeling DatasetAngel Daza, Anette FrankEMNLP 2020 · 被引用 21 次
- OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information ExtractionKeshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Mausam 等EMNLP 2020 · 被引用 13 次
- IMoJIE: Iterative Memory-Based Joint Open Information ExtractionKeshav Kolluru, Samarth Aggarwal, Vipul Rathore, Mausam 等ACL 2020 · 被引用 5 次
相关 Paper
- DetIE: Multilingual Open Information Extraction Inspired by Object DetectionMichael Vasilkovsky, Anton Alekseev, Valentin Malykh, Ilya Shenbin 等AAAI 2022 · 被引用 24 次
- WebIE: Faithful and Robust Information Extraction on the WebChenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos 等ACL 2023 · 被引用 3 次
- LOREM: Language-consistent Open Relation Extraction from Unstructured TextTom Harting, Sepideh Mesbah, Christoph LofiWWW 2020 · 被引用 4 次
- Syntactically Rich Discriminative Training: An Effective Method for Open Information ExtractionFrank Mtumbuka, Thomas LukasiewiczEMNLP 2022 · 被引用 1 次
- Abstractive Open Information ExtractionKevin Pei, Ishan Jindal, Kevin Chen-Chuan ChangEMNLP 2023
