Semantics Altering Modifications for Evaluating Comprehension in Machine Reading
Viktor Schlegel, Goran Nenadic, Riza Batista-Navarro
摘要
Advances in NLP have yielded impressive results for the task of machine reading comprehension (MRC), with approaches having been reported to achieve performance comparable to that of humans. In this paper, we investigate whether state-of-the-art MRC models are able to correctly process Semantics Altering Modifications (SAM): linguistically-motivated phenomena that alter the semantics of a sentence while preserving most of its lexical surface form. We present a method to automatically generate and align challenge sets featuring original and altered examples. We further propose a novel evaluation methodology to correctly assess the capability of MRC systems to process these examples independent of the data they were optimised on, by discounting for effects introduced by domain shift. In a large-scale empirical study, we apply the methodology in order to evaluate extractive MRC models with regard to their capability to correctly process SAM-enriched data. We comprehensively cover 12 different state-of-the-art neural architecture configurations and four training datasets and find that -- despite their well-known remarkable performance -- optimised models consistently struggle to correctly process semantically altered data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringWang Zhu, Jesse Thomason, Robin JiaEMNLP 2023 · 被引用 2 次
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 被引用 2 次
它引用的顶会 Paper3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
相关 Paper
- Validation on machine reading comprehension software without annotated labels: a property-based methodSongqiang Chen, Shuo Jin, Xiaoyuan XieFSE 2021 · 被引用 27 次
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 被引用 92 次
- Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language ModelsYang Liu, Hongming Li, Melissa Xiaohui Qin, Chao Huang 等ACL 2026
- A Robust Adversarial Training Approach to Machine Reading ComprehensionKai Liu, Xin Liu, An Yang, Jing Liu 等AAAI 2020 · 被引用 54 次
- Robust Domain Adaptation for Machine Reading ComprehensionLiang Jiang, Zhenyu Huang, Jia Liu, Zujie Wen 等AAAI 2023 · 被引用 1 次
