Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word Game
Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan
Abstract
The New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts. We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice human players. Our results show that even the bestperforming LLM, Claude 3.5 Sonnet, which has otherwise shown impressive reasoning abilities on a wide variety of benchmarks, can only fully solve 18% of the games. Novice and expert players perform better than Claude 3.5 Sonnet, with expert human players significantly outperforming it. We create a taxonomy of the knowledge types required to successfully cluster and categorize words in the Connections game. We find that while LLMs perform relatively well on categorizing words based on semantic relations they struggle with other types of knowledge such as Encyclopedic Knowledge, Multiword Expressions or knowledge that combines both Word Form and Meaning. Our results establish the New York Times Connections game as a challenging benchmark for evaluating abstract reasoning capabilities in AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bda45070-8e7c-4ab5-9cd3-9a38f7644586Cited by top-tier papers3
- How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in ItalianAndrea Pedrotti, Giulia Rambelli, Caterina Villani, Marianna BolognesiACL 2025 · 2 citations
- Think Globally, Group Locally: Evaluating LLMs Using Multi-Lingual Word Grouping GamesCésar Guerra-Solano, Zhuochun Li, Xiang Lorraine LiEMNLP 2025
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps UsersNishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant et al.EMNLP 2025
Builds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- SmartPlay : A Benchmark for LLMs as Intelligent AgentsYue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi LiICLR 2024 · 121 citations
- Abstract Visual Reasoning with Tangram ShapesAnya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr et al.EMNLP 2022 · 18 citations
- Automated Crossword SolvingEric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang et al.ACL 2022 · 17 citations
Related papers
- Down and Across: Introducing Crossword-Solving as a New NLP BenchmarkSaurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, Anna RumshiskyACL 2022 · 5 citations
- FCMR: Robust Evaluation of Financial Cross-Modal Multi-Hop ReasoningSeunghee Kim, Changhyeon Kim, Taeuk KimACL 2025
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 9 citations
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language ModelsSophia Simeng Han, Howard Dai, Stephen Xia, Grant Zhang et al.NeurIPS 2025 · 2 citations
- How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human ComparisonJiayin Wang, Zhiqiang Guo, Weizhi Ma, Min ZhangEMNLP 2025
