Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
Neh Majmudar, Elena Filatova
Abstract
In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles.We focus on puzzles used in Linguistic Olympiads for high school students.We first extend the existing benchmark for the task of solving linguistic puzzles.We explore the use of Large Language Models (LLMs), including recent state-of-the-art models such as OpenAI's o1, for solving linguistic puzzles, analyzing their performance across various linguistic topics.We demonstrate that LLMs outperform humans on most puzzles types, except for those centered on writing systems, and for the understudied languages.We use the insights from puzzle-solving experiments to direct the novel task of puzzle generation.We believe that automating puzzle generation, even for relatively simple puzzles, holds promise for expanding interest in linguistics and introducing the field to a broader audience.This finding highlights the importance of linguistic puzzle generation as a research task: such puzzles can not only promote linguistics but also support the dissemination of knowledge about rare and understudied languages.C.1.9GPT-4o, Few-shot, Spanish C.2.8 OpenAI's o1, Few-shot, Gujarati
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75e88873-ba68-44ae-bbed-6842935d3897Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksBingsheng Yao, Dakuo Wang, Tongshuang Wu, Zheng Zhang et al.ACL 2022 · 58 citations
- PathQG: Neural Question Generation from FactsSiyuan Wang, Zhongyu Wei, Zhihao Fan, Zengfeng Huang et al.EMNLP 2020 · 18 citations
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 9 citations
Related papers
- PuzzLing Machines: A Challenge on Learning From Small DataGözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna GurevychACL 2020
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 6 citations
- Logic.py: Bridging the Gap between LLMs and Constraint SolversPascal Kesseli, Peter W. O'Hearn, Ricardo Silveira CabralNeurIPS 2025 · 10 citations
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai et al.ICLR 2025 · 3 citations
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu et al.ICML 2026 · 1 citation
