BRAINTEASER: Lateral Thinking Puzzles for Large Language Models
Yifan Jiang, Filip Ilievski, Kaixin Ma, Zhivar Sourati
摘要
The success of language models has inspired the NLP community to attend to tasks that require implicit and complex reasoning, relying on human-like commonsense mechanisms. While such vertical thinking tasks have been relatively popular, lateral thinking puzzles have received little attention. To bridge this gap, we devise BRAINTEASER: a multiple-choice Question Answering task designed to test the model's ability to exhibit lateral thinking and defy default commonsense associations. We design a three-step procedure for creating the first lateral thinking benchmark, consisting of data collection, distractor generation, and generation of reconstruction examples, leading to 1,100 puzzles with high-quality annotations. To assess the consistency of lateral reasoning by models, we enrich BRAINTEASER based on a semantic and contextual reconstruction of its questions. Our experiments with state-of-the-art instruction- and commonsense language models reveal a significant gap between human and model performance, which is further widened when consistency across reconstruction formats is considered. We make all of our code and data available to stimulate work on developing and evaluating lateral thinking models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 被引用 9 次
- COLUMBUS: Evaluating COgnitive Lateral Understanding Through Multiple-Choice reBUSesKoen Kraaijveld, Yifan Jiang, Kaixin Ma, Filip IlievskiAAAI 2025 · 被引用 8 次
- MP: Endowing Large Language Models with Lateral ThinkingTian Bai, Yongwang Cao, Yan Ge, Haitao YuAAAI 2025 · 被引用 4 次
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang 等NeurIPS 2025 · 被引用 2 次
- Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social InteractionsMaojia Song, Pala Tej Deep, Ruiwen Zhou, Weisheng Jin 等ICLR 2026
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation PuzzlesQi Chen, Bowen Zhang, Gang Wang, Qi WuNeurIPS 2024 · 被引用 13 次
- Down and Across: Introducing Crossword-Solving as a New NLP BenchmarkSaurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, Anna RumshiskyACL 2022 · 被引用 5 次
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu 等CVPR 2026 · 被引用 7 次
- VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain KnowledgeYueqi Song, Tianyue Ou, Yibo Kong, Zecheng Li 等ICML 2026 · 被引用 44 次
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 被引用 38 次
