PuzzLing Machines: A Challenge on Learning From Small Data
Gözde Gül Sahin, Yova Kementchedjhieva, Phillip Rust, Iryna Gurevych
Abstract
Deep neural models have repeatedly proved excellent at memorizing surface patterns from large datasets for various ML and NLP benchmarks. They struggle to achieve human-like thinking, however, because they lack the skill of iterative reasoning upon knowledge. To expose this problem in a new light, we introduce a challenge on learning from small data, PuzzLing Machines, which consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students. These puzzles are carefully designed to contain only the minimal amount of parallel text necessary to deduce the form of unseen expressions. Solving them does not require external information (e.g., knowledge bases, visual signals) or linguistic expertise, but meta-linguistic awareness and deductive skills. Our challenge contains around 100 puzzles covering a wide range of linguistic phenomena from 81 languages. We show that both simple statistical algorithms and state-of-the-art deep neural models perform inadequately on this challenge, as expected. We hope that this benchmark, available at https://ukplab.github.io/ PuzzLing-Machines/ , inspires further efforts towards a new paradigm in NLP-one that is grounded in human-like reasoning and understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e3cf892-94d9-4963-829b-5adfafb93027Cited by top-tier papers4
- LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic ObfuscationJude Khouja, Lingyi Yang, Karolina Korgul, Simeon Hellsten et al.ICLR 2026 · 4 citations
- Can LLMs Generate and Solve Linguistic Olympiad Puzzles?Neh Majmudar, Elena FilatovaEMNLP 2025 · 2 citations
- Glider: Global and Local Instruction-Driven Expert RouterPingzhi Li, Prateek Yadav, Jaehong Yoon, Jie Peng et al.EMNLP 2025 · 1 citation
- Explicit Learning and the LLM in Machine TranslationMalik Marmonier, Rachel Bawden, Benoît SagotEMNLP 2025
Related papers
- Are Deep Neural Networks SMARTer Than Second Graders?Anoop Cherian, Kuan-Chuan Peng, Suhas Lohit, Kevin A. Smith et al.CVPR 2023
- Learning Compositional Rules via Neural Program SynthesisMaxwell I. Nye, Armando Solar-Lezama, Josh Tenenbaum, Brenden M. LakeNeurIPS 2020 · 120 citations
- Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning ProcessTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 3 citations
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMsZhongshen Zeng, Yinhong Liu, Yingjia Wan, Jingyao Li et al.NeurIPS 2024 · 51 citations
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesJiangjie Chen, Qianyu He, Siyu Yuan, Aili Chen et al.NeurIPS 2025 · 60 citations
