BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies?
Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, José Camacho-Collados
摘要
Analogies play a central role in human commonsense reasoning. The ability to recognize analogies such as "eye is to seeing what ear is to hearing", sometimes referred to as analogical proportions, shape how we structure knowledge and understand language. Surprisingly, however, the task of identifying such analogies has not yet received much attention in the language model era. In this paper, we analyze the capabilities of transformer-based language models on this unsupervised task, using benchmarks obtained from educational settings, as well as more commonly used datasets. We find that off-the-shelf language models can identify analogies to a certain extent, but struggle with abstract and complex relations, and results are highly sensitive to model architecture and hyperparameters. Overall the best results were obtained with GPT-2 and RoBERTa, while configurations using BERT were not able to outperform word embedding models. Our results raise important questions for future work about how, and to what extent, pre-trained language models capture knowledge about abstract semantic relations. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsMatthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille 等ICCV 2023 · 被引用 51 次
- Multimodal Analogical Reasoning over Knowledge GraphsNingyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang 等ICLR 2023 · 被引用 10 次
- Emergent Analogical Reasoning in TransformersGouki Minegishi, Jingyuan Feng, Hiroki Furuta, Takeshi Kojima 等ICML 2026 · 被引用 4 次
- Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performanceMolly R. Petersen, Lonneke van der PlasEMNLP 2023 · 被引用 3 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 被引用 183 次
- Making Pre-trained Language Models Better Few-shot LearnersTianyu Gao, Adam Fisch, Danqi ChenACL 2021
相关 Paper
- AnaloBench: Benchmarking the Identification of Abstract and Long-context AnalogiesXiao Ye, Andrew Wang, Jacob Choi, Yining Lu 等EMNLP 2024 · 被引用 3 次
- In-Context Analogical Reasoning with Pre-Trained Language ModelsXiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce ChaiACL 2023 · 被引用 13 次
- KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal ModelsEunice Yiu, Maan Qraitem, Anisa Noor Majhi, Charlie Wong 等ICLR 2025
- Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense ReasoningRuben Branco, António Branco, João António Rodrigues, João Ricardo SilvaEMNLP 2021 · 被引用 29 次
- Can Pre-trained Language Models Interpret Similes as Smart as Human?Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie 等ACL 2022
