WebIE: Faithful and Robust Information Extraction on the Web
Chenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos, Andrea Pierleoni
摘要
Extracting structured and grounded fact triples from raw text is a fundamental task in Information Extraction (IE). Existing IE datasets are typically collected from Wikipedia articles, using hyperlinks to link entities to the Wikidata knowledge base. However, models trained only on Wikipedia have limitations when applied to web domains, which often contain noisy text or text that does not have any factual information. We present WEBIE, the first large-scale, entity-linked closed IE dataset consisting of 1.6M sentences automatically collected from the English Common Crawl corpus. WEBIE also includes negative examples, i.e. sentences without fact triples, to better reflect the data on the web. We annotate ∼21K triples from WEBIE through crowdsourcing and introduce mWEBIE, a translation of the annotated set in four other languages: French, Spanish, Portuguese, and Hindi. We evaluate the in-domain, out-of-domain, and zero-shot cross-lingual performance of generative IE models and find models trained on WEBIE show better generalisability. We also propose three training strategies that use entity linking as an auxiliary task. Our experiments show that adding Entity-Linking objectives improves the faithfulness of our generative IE models 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda 等EMNLP 2020 · 被引用 562 次
- Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation ExtractionTapas Nayak, Hwee Tou NgAAAI 2020 · 被引用 272 次
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 被引用 200 次
- Re-TACRED: Addressing Shortcomings of the TACRED DatasetGeorge Stoica, Emmanouil Antonios Platanios, Barnabás PóczosAAAI 2021 · 被引用 146 次
相关 Paper
- Entity Linking in 100 LanguagesJan A. Botha, Zifei Shan, Daniel GillickEMNLP 2020
- SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific DocumentsQi Zhang, Zhijia Chen, Huitong Pan, Cornelia Caragea 等EMNLP 2024 · 被引用 7 次
- An Analysis of Multilingual FActScoreVu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt 等EMNLP 2024 · 被引用 2 次
- RAED: Retrieval-Augmented Entity Description Generation for Emerging Entity Linking and DisambiguationKarim Ghonim, Pere-Lluís Huguet Cabot, Riccardo Orlando, Roberto NavigliEMNLP 2025
- Increasing Coverage and Precision of Textual Information in Multilingual Knowledge GraphsSimone Conia, Min Li, Daniel Lee, Umar Farooq Minhas 等EMNLP 2023 · 被引用 3 次
