Can LLMs Generate and Solve Linguistic Olympiad Puzzles?
Neh Majmudar, Elena Filatova
摘要
In this paper, we introduce a combination of novel and exciting tasks: the solution and generation of linguistic puzzles.We focus on puzzles used in Linguistic Olympiads for high school students.We first extend the existing benchmark for the task of solving linguistic puzzles.We explore the use of Large Language Models (LLMs), including recent state-of-the-art models such as OpenAI's o1, for solving linguistic puzzles, analyzing their performance across various linguistic topics.We demonstrate that LLMs outperform humans on most puzzles types, except for those centered on writing systems, and for the understudied languages.We use the insights from puzzle-solving experiments to direct the novel task of puzzle generation.We believe that automating puzzle generation, even for relatively simple puzzles, holds promise for expanding interest in linguistics and introducing the field to a broader audience.This finding highlights the importance of linguistic puzzle generation as a research task: such puzzles can not only promote linguistics but also support the dissemination of knowledge about rare and understudied languages.C.1.9GPT-4o, Few-shot, Spanish C.2.8 OpenAI's o1, Few-shot, Gujarati
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksBingsheng Yao, Dakuo Wang, Tongshuang Wu, Zheng Zhang 等ACL 2022 · 被引用 58 次
- PathQG: Neural Question Generation from FactsSiyuan Wang, Zhongyu Wei, Zhihao Fan, Zengfeng Huang 等EMNLP 2020 · 被引用 18 次
- Puzzle Solving using Reasoning of Large Language Models: A SurveyPanagiotis Giadikiaroglou, Maria Lymperaiou, Giorgos Filandrianos, Giorgos StamouEMNLP 2024 · 被引用 9 次
相关 Paper
- PuzzLing Machines: A Challenge on Learning From Small DataGözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna GurevychACL 2020
- "A good pun is its own reword": Can Large Language Models Understand Puns?Zhijun Xu, Siyu Yuan, Lingjie Chen, Deqing YangEMNLP 2024 · 被引用 6 次
- Logic.py: Bridging the Gap between LLMs and Constraint SolversPascal Kesseli, Peter W. O'Hearn, Ricardo Silveira CabralNeurIPS 2025 · 被引用 10 次
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai 等ICLR 2025 · 被引用 3 次
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu 等ICML 2026 · 被引用 1 次
