Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game
Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan
摘要
The New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts. We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice human players. Our results show that even the bestperforming LLM, Claude 3.5 Sonnet, which has otherwise shown impressive reasoning abilities on a wide variety of benchmarks, can only fully solve 18% of the games. Novice and expert players perform better than Claude 3.5 Sonnet, with expert human players significantly outperforming it. We create a taxonomy of the knowledge types required to successfully cluster and categorize words in the Connections game. We find that while LLMs perform relatively well on categorizing words based on semantic relations they struggle with other types of knowledge such as Encyclopedic Knowledge, Multiword Expressions or knowledge that combines both Word Form and Meaning. Our results establish the New York Times Connections game as a challenging benchmark for evaluating abstract reasoning capabilities in AI systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in ItalianAndrea Pedrotti, Giulia Rambelli, Caterina Villani, Marianna BolognesiACL 2025 · 被引用 2 次
- Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping GamesCésar Guerra-Solano, Zhuochun Li, Xiang Lorraine LiEMNLP 2025
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps UsersNishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant 等EMNLP 2025
它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 被引用 121 次
- Abstract Visual Reasoning with Tangram ShapesAnya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr 等EMNLP 2022 · 被引用 18 次
- Automated Crossword SolvingEric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang 等ACL 2022 · 被引用 17 次
相关 Paper
- Down and Across: Introducing Crossword-Solving as a New NLP BenchmarkSaurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, Anna RumshiskyACL 2022 · 被引用 5 次
- FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop ReasoningSeunghee Kim, Changhyeon Kim, Taeuk KimACL 2025
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 被引用 9 次
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang 等NeurIPS 2025 · 被引用 2 次
- How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human ComparisonJiayin Wang, Zhiqiang Guo, Weizhi Ma, Min ZhangEMNLP 2025
