The Two-Hump Problem: Bridging the Difficulty Gap in Mathematical Reinforcement Learning
Lucas Fagan, Michele Tarquini, Ali Shehper, Maksymilian Manko, Angus Gruen, Coco Huang, Giorgi Butbaia, Davide Passaro, Sergei Gukov
Abstract
Mathematical search problems present a unique challenge for Reinforcement Learning (RL) due to vast search spaces and sparse rewards. In previous works, the Andrews-Curtis (AC) conjecture was established as an illustrative example of such problems. In this work, we identify a critical structural barrier in the AC landscape: a "Two Hump" distribution, where problem instances are either trivially solvable or effectively impossible, with a scarcity of intermediate "hard-but-solvable" instances required for effective learning. We tackle this challenge through two primary avenues: novel data generation techniques to populate the difficulty gap, and significant algorithmic enhancements including the introduction of supermoves and Transformer-based architectures. We demonstrate substantial performance improvements over previous baselines, and release new comprehensive benchmark datasets including AC-19 (125,192 AC-trivial presentations of varying difficulty with length at most 19) and AC-1M (1,136,154 hard AC-trivial presentations of length at most 30), the first large-scale, publicly available datasets of this kind.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4805603-c19e-4280-a83d-c0e582aab6c9Builds on3
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Formal Mathematics Statement Curriculum LearningStanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys et al.ICLR 2023 · 24 citations
- What makes math problems hard for reinforcement learning: A case studyAli Shehper, Anibal M. Medina-Mardones, Lucas Fagan, Bartlomiej Lewandowski et al.NeurIPS 2025 · 15 citations
Related papers
- D²Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement LearningRu Zhang, Renda Li, Ziyu Ma, Weijie Qiu et al.ICML 2026
- Machine Learning meets Algebraic Combinatorics: A Suite of Datasets Capturing Research-level Conjecturing Ability in Pure MathematicsHerman Chau, Helen Jenne, Davis Brown, Jesse He et al.ICML 2025
- Self-Supervised Transformers as Iterative Solution Improvers for Constraint SatisfactionYudong Xu, Wenhao Li, Scott Sanner, Elias Boutros KhalilICML 2025
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja et al.ICML 2026 · 16 citations
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing ReasoningZhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu et al.ICLR 2026 · 271 citations
