Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction
Ji Qi, Chuchun Zhang, Xiaozhi Wang, Kaisheng Zeng, Jifan Yu, Jinxin Liu, Lei Hou, Juanzi Li, Xu Bin
Abstract
The robustness to distribution changes ensures that NLP models can be successfully applied in the realistic world, especially for information extraction tasks. However, most prior evaluation benchmarks have been devoted to validating pairwise matching correctness, ignoring the crucial validation of robustness. In this paper, we present the first benchmark that simulates the evaluation of open information extraction models in the real world, where the syntactic and expressive distributions under the same knowledge meaning may drift variously. We design and annotate a large-scale testbed in which each example is a knowledge-invariant clique that consists of sentences with structured knowledge of the same meaning but with different syntactic and expressive forms. By further elaborating the robustness metric, a model is judged to be robust if its performance is consistently accurate on the overall cliques. We perform experiments on typical models published in the last decade as well as a representative large language model, and the results show that the existing successful models exhibit a frustrating degradation, with a maximum drop of 23.43 F 1 score. Our resources and code are available at https://github.com/qijimrc/ROBUST .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a63d0a19-052b-48a8-8984-9122f50fac42Cited by top-tier papers3
- Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language ModelsQingkai Min, Qipeng Guo, Xiangkun Hu, Songfang Huang et al.ACL 2024 · 11 citations
- ADELIE: Aligning Large Language Models on Information ExtractionYunjia Qi, Hao Peng, Xiaozhi Wang, Bin Xu et al.EMNLP 2024 · 8 citations
- Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction SystemsItalo Luis da Silva, Hanqi Yan, Lin Gui, Yulan HeEMNLP 2024
Builds on9
- Span Model for Open Information Extraction on Accurate CorpusJunlang Zhan, Hai ZhaoAAAI 2020 · 90 citations
- Generalizing Natural Language Analysis through Span-relation RepresentationsZhengbao Jiang, Wei Xu, Jun Araki, Graham NeubigACL 2020 · 58 citations
- AESOP: Paraphrase Generation with Adaptive Syntactic ControlJiao Sun, Xuezhe Ma, Nanyun PengEMNLP 2021 · 47 citations
- Knowledge Infused DecodingRuibo Liu, Guoqing Zheng, Shashank Gupta, Radhika Gaonkar et al.ICLR 2022 · 18 citations
- OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information ExtractionKeshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Mausam et al.EMNLP 2020 · 13 citations
Related papers
- OODREB: Benchmarking State-of-the-Art Methods for Out-Of-Distribution Generalization on Relation ExtractionHaotian Chen, Houjing Guo, Bingsheng Chen, Xiangdong ZhouWWW 2024 · 1 citation
- Towards Robust Universal Information Extraction: Dataset, Evaluation, and SolutionJizhao Zhu, Akang Shi, Zixuan Li, Long Bai et al.ACL 2025
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed InputsZhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim et al.EMNLP 2025
- The Effect of Natural Distribution Shift on Question Answering ModelsJohn Miller, Karl Krauth, Benjamin Recht, Ludwig SchmidtICML 2020 · 158 citations
