The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources
Akshatha Arodi, Martin Pömsl, Kaheer Suleman, Adam Trischler, Alexandra Olteanu, Jackie Chi Kit Cheung
摘要
Many state-of-the-art natural language understanding (NLU) models are based on pretrained neural language models. These models often make inferences using information from multiple sources. An important class of such inferences are those that require both background knowledge, presumably contained in a model's pretrained parameters, and instance-specific information that is supplied at inference time. However, the integration and reasoning abilities of NLU models in the presence of multiple knowledge sources have been largely understudied. In this work, we propose a test suite of coreference resolution subtasks that require reasoning over multiple facts. These subtasks differ in terms of which knowledge sources contain the relevant facts. We also introduce subtasks where knowledge is present only at inference time using fictional knowledge. We evaluate state-of-the-art coreference resolution models on our dataset. Our results indicate that several models struggle to reason on-thefly over knowledge observed both at pretrain time and at inference time. However, with taskspecific training, a subset of models demonstrates the ability to integrate certain knowledge types from multiple sources. Still, even the best performing models seem to have difficulties with reliably integrating knowledge presented only at inference time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive MimickingTimo Schick, Hinrich SchützeAAAI 2020 · 被引用 106 次
- RICA: Evaluating Robust Inference Capabilities Based on Commonsense AxiomsPei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin 等EMNLP 2021 · 被引用 28 次
- Toward Gender-Inclusive Coreference ResolutionYang Trista Cao, Hal Daumé IIIACL 2020 · 被引用 20 次
- Entity-Based Knowledge Conflicts in Question AnsweringShayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh 等EMNLP 2021 · 被引用 3 次
相关 Paper
- Attending to Entities for Better Text UnderstandingPengxiang Cheng, Katrin ErkAAAI 2020 · 被引用 42 次
- Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut 等ACL 2020 · 被引用 168 次
- Probing Linguistic Information for Logical Inference in Pre-trained Language ModelsZeming Chen, Qiyue GaoAAAI 2022 · 被引用 11 次
- Seq2seq is All You Need for Coreference ResolutionWenzheng Zhang, Sam Wiseman, Karl StratosEMNLP 2023 · 被引用 5 次
- Constrained Multi-Task Learning for Bridging ResolutionHideo Kobayashi, Yufang Hou, Vincent NgACL 2022
