Alignment-Augmented Consistent Translation for Multilingual Open Information Extraction
Keshav Kolluru, Muqeeth Mohammed, Shubham Mittal, Soumen Chakrabarti, Mausam
Abstract
Progress with supervised Open Information Extraction (OpenIE) has been primarily limited to English due to the scarcity of training data in other languages. In this paper, we explore techniques to automatically convert English text for training OpenIE systems in other languages. We introduce the Alignment-Augmented Constrained Translation (AACTrans) model to translate English sentences and their corresponding extractions consistently with each other — with no changes to vocabulary or semantic meaning which may result from independent translations. Using the data generated with AACTrans, we train a novel two-stage generative OpenIE model, which we call Gen2OIE, that outputs for each sentence: 1) relations in the first stage and 2) all extractions containing the relation in the second stage. Gen2OIE increases relation coverage using a training data transformation technique that is generalizable to multiple languages, in contrast to existing models that use an English-specific training loss. Evaluations on 5 languages — Spanish, Portuguese, Chinese, Hindi and Telugu — show that the Gen2OIE with AACTrans data outperforms prior systems by a margin of 6-25% in F1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05550960-ca7e-41c7-a509-4f69344a2acfCited by top-tier papers7
- Multilingual Relation Classification via Efficient and Effective PromptingYuxuan Chen, David Harbecke, Leonhard HennigEMNLP 2022 · 13 citations
- Learning to Extract Structured Entities Using Language ModelsHaolun Wu, Ye Yuan, Liana Mikaelyan, Alexander Meulemans et al.EMNLP 2024 · 5 citations
- Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social MediaShubham Mittal, Megha Sundriyal, Preslav NakovEMNLP 2023 · 4 citations
- MultiTACRED: A Multilingual Version of the TAC Relation Extraction DatasetLeonhard Hennig, Philippe Thomas, Sebastian MöllerACL 2023 · 4 citations
- "Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun CompoundsKeshav Kolluru, Gabriel Stanovsky, MausamEMNLP 2022 · 1 citation
Builds on4
- ParaCrawl: Web-Scale Acquisition of Parallel CorporaMarta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield et al.ACL 2020 · 132 citations
- X-SRL: A Parallel Cross-Lingual Semantic Role Labeling DatasetAngel Daza, Anette FrankEMNLP 2020 · 21 citations
- OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information ExtractionKeshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Mausam et al.EMNLP 2020 · 13 citations
- IMoJIE: Iterative Memory-Based Joint Open Information ExtractionKeshav Kolluru, Samarth Aggarwal, Vipul Rathore, Mausam et al.ACL 2020 · 5 citations
Related papers
- DetIE: Multilingual Open Information Extraction Inspired by Object DetectionMichael Vasilkovsky, Anton Alekseev, Valentin Malykh, Ilya Shenbin et al.AAAI 2022 · 24 citations
- WebIE: Faithful and Robust Information Extraction on the WebChenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos et al.ACL 2023 · 3 citations
- LOREM: Language-consistent Open Relation Extraction from Unstructured TextTom Harting, Sepideh Mesbah, Christoph LofiWWW 2020 · 4 citations
- Syntactically Rich Discriminative Training: An Effective Method for Open Information ExtractionFrank Mtumbuka, Thomas LukasiewiczEMNLP 2022 · 1 citation
- Abstractive Open Information ExtractionKevin Pei, Ishan Jindal, Kevin Chen-Chuan ChangEMNLP 2023
