What's the Best Place for an AI Conference, Vancouver or _______: Why Completing Comparative Questions is Difficult
Avishai Zagoury, Einat Minkov, Idan Szpektor, William W. Cohen
摘要
Although large neural language models (LMs) like BERT can be finetuned to yield state-of-the-art results on many NLP tasks, it is often unclear what these models actually learn. Here we study using such LMs to fill in entities in human-authored comparative questions, like ``Which country is older, India or _____?''---i.e., we study the ability of neural LMs to ask (not answer) reasonable questions. We show that accuracy in this fill-in-the-blank task is well-correlated with human judgements of whether a question is reasonable, and that these models can be trained to achieve nearly human-level performance in completing comparative questions in three different subdomains. However, analysis shows that what they learn fails to model any sort of broad notion of which entities are semantically comparable or similar---instead the trained models are very domain-specific, and performance is highly correlated with co-occurrences between specific entities observed in the training set. This is true both for models that are pretrained on general text corpora, as well as models trained on a large corpus of comparison questions. Our study thus reinforces recent results on the difficulty of making claims about a deep model's world knowledge or linguistic competence based on performance on specific benchmark problems. We make our evaluation datasets publicly available to foster future research on complex understanding and reasoning in such models at standards of human interaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Life after BERT: What do Other Muppets Understand about Language?Vladislav Lialin, Kevin Zhao, Namrata Shivagunde, Anna RumshiskyACL 2022 · 被引用 7 次
- Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word DistributionsHaw-Shiuan Chang, Andrew McCallumACL 2022
- Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language UnderstandingMorris Alper, Michael Fiman, Hadar Averbuch-ElorCVPR 2023
它引用的顶会 Paper4
- Probing Natural Language Inference Models through Semantic FragmentsKyle Richardson, Hai Hu, Lawrence S. Moss, Ashish SabharwalAAAI 2020 · 被引用 152 次
- Entities as Experts: Sparse Memory Access with Entity SupervisionThibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi 等EMNLP 2020 · 被引用 39 次
- How does BERT's attention change when you fine-tune? An analysis methodology and a case study in negation scopeYiyun Zhao, Steven BethardACL 2020 · 被引用 35 次
- Dynamic Composition for Conversational Domain ExplorationIdan Szpektor, Deborah Cohen, Gal Elidan, Michael Fink 等WWW 2020 · 被引用 11 次
相关 Paper
- IELM: An Open Information Extraction Benchmark for Pre-Trained Language ModelsChenguang Wang, Xiao Liu, Dawn SongEMNLP 2022 · 被引用 3 次
- Pre-training Language Models for Comparative ReasoningMengxia Yu, Zhihan Zhang, Wenhao Yu, Meng JiangEMNLP 2023 · 被引用 1 次
- Can LMs Learn New Entities from Descriptions? Challenges in Propagating Injected KnowledgeYasumasa Onoe, Michael J. Q. Zhang, Shankar Padmanabhan, Greg Durrett 等ACL 2023 · 被引用 24 次
- X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language ModelsZhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding 等EMNLP 2020 · 被引用 81 次
- Can Pre-trained Language Models Interpret Similes as Smart as Human?Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie 等ACL 2022
