I am a Strange Dataset: Metalinguistic Tests for Language Models
Tristan Thrush, Jared Moore, Miguel Monares, Christopher Potts, Douwe Kiela
摘要
Statements involving metalinguistic selfreference ("This paper has six sections.") are prevalent in many domains. Can large language models (LLMs) handle such language? In this paper, we present "I am a Strange Dataset", a new dataset for addressing this question. There are two subtasks: generation and verification. In generation, models continue statements like "The penultimate word in this sentence is" (where a correct continuation is "is"). In verification, models judge the truth of statements like "The penultimate word in this sentence is sentence." (false). We also provide minimally different metalinguistic non-self-reference examples to complement the main dataset by probing for whether models can handle metalinguistic language at all. The dataset is hand-crafted by experts and validated by non-expert annotators. We test a variety of open-source LLMs (7B to 70B parameters) as well as closed-source LLMs through APIs. All models perform close to chance across both subtasks and even on the non-self-referential metalinguistic control data, though we find some steady improvement with model scale. GPT 4 is the only model to consistently do significantly better than chance, and it is still only in the 60% range, while our untrained human annotators score well in the 89-93% range. one paper you read today is bound to contain "In 044 this paper" (Anonymous, 2024). 045 In this paper, we focus on metalinguistic self-046 reference, the complex kind of self-reference in 047 which language is used to make claims about it-048 self, as in "This sentence has five words" and "This 049 paper has six sections". 1 Using such language in-050 volves reasoning about metalinguistic properties 051 (counting words, naming parts of speech, etc.) and 052 resolving self-reference. Humans generally have 053 no trouble with such language, and may even enjoy 054 its playful and sometimes paradoxical nature (Hof-055
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- Winoground: Probing Vision and Language Models for Visio-Linguistic CompositionalityTristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh 等CVPR 2022 · 被引用 179 次
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for LittleKoustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau 等EMNLP 2021 · 被引用 177 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
相关 Paper
- Metaphor Understanding Challenge Dataset for LLMsXiaoyu Tong, Rochelle Choenni, Martha Lewis, Ekaterina ShutovaACL 2024
- M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMsTianlong Zheng, Yating Yang, Rui Dong, Bo Ma 等AAAI 2026
- This is not a Dataset: A Large Negation Benchmark to Challenge Large Language ModelsIker García-Ferrero, Begoña Altuna, Javier Álvez, Itziar Gonzalez-Dios 等EMNLP 2023 · 被引用 8 次
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementWenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan 等ACL 2024
- ANAH: Analytical Annotation of Hallucinations in Large Language ModelsZiwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu 等ACL 2024 · 被引用 8 次
