Down and Across: Introducing Crossword-Solving as a New NLP Benchmark
Saurabh Kulshreshtha, Olga Kovaleva, Namrata Shivagunde, Anna Rumshisky
Abstract
Solving crossword puzzles requires diverse reasoning capabilities, access to a vast amount of knowledge about language and the world, and the ability to satisfy the constraints imposed by the structure of the puzzle. In this work, we introduce solving crossword puzzles as a new natural language understanding task. We release the specification of a corpus of crossword puzzles collected from the New York Times daily crossword spanning 25 years and comprised of a total of around nine thousand puzzles. These puzzles include a diverse set of clues: historic, factual, word meaning, synonyms/antonyms, fillin-the-blank, abbreviations, prefixes/suffixes, wordplay, and cross-lingual, as well as clues that depend on the answers to other clues. We separately release the clue-answer pairs from these puzzles as an open-domain question answering dataset containing over half a million unique clue-answer pairs. For the question answering task, our baselines include several sequence-to-sequence and retrieval-based generative models. We also introduce a nonparametric constraint satisfaction baseline for solving the entire crossword puzzle. Finally, we propose an evaluation framework which consists of several complementary performance metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62cab57a-d840-4fc9-acd7-e1916a1f8e7fCited by top-tier papers1
Ask how each one uses itBuilds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Assessing the Benchmarking Capacity of Machine Reading Comprehension DatasetsSaku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko AizawaAAAI 2020 · 92 citations
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 54 citations
Related papers
- Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLPJosh Rozner, Christopher Potts, Kyle MahowaldNeurIPS 2021 · 27 citations
- Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word GamePrisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf et al.EMNLP 2024 · 5 citations
- Automated Crossword SolvingEric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang et al.ACL 2022 · 17 citations
- A Reasoning-Based Approach to Cryptic Crossword Clue SolvingMartin Andrews, Sam WitteveenICML 2025
- BRAINTEASER: Lateral Thinking Puzzles for Large Language ModelsYifan Jiang, Filip Ilievski, Kaixin Ma, Zhivar SouratiEMNLP 2023 · 6 citations
