e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Amrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer, Ian Wu, Virginia Smith, Max Simchowitz, Aviral Kumar
Abstract
Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.e., improvement in performance on hard problems as LLMs keep "thinking" for longer, beyond the maximum token budget they were trained on). Surprisingly, we find that most existing reasoning models do not extrapolate well. We show that one way to enable extrapolation is by training the LLM to perform in-context exploration: training the LLM to effectively spend its test time budget by chaining operations (such as generation, verification, refinement, etc.), or testing multiple hypotheses before it commits to an answer. To enable in-context exploration, we identify three key ingredients as part of our recipe e3: (1) chaining skills that the base LLM has asymmetric competence in, e.g., chaining verification (easy) with generation (hard), as a way to implement in-context search; (2) leveraging "negative" gradients from incorrect traces to amplify exploration during RL, resulting in longer search traces that chains additional asymmetries; and (3) coupling task difficulty with training token budget during training via a specifically-designed curriculum to structure in-context exploration. Our recipe e3 produces the best known 1.7B model according to AIME'25 and HMMT'25 scores, and extrapolates to 2x the training token budget. Our e3-1.7B model not only attains high pass@1 scores, but also improves pass@k over the base model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4516827-a8e6-4843-a4ae-b3995069dcb7Cited by top-tier papers3
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic TasksHao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar et al.CVPR 2026
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsTejas Krishnan, Sumeet Motwani, Charles London, Suhaas Bhat et al.ICML 2026
- OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesSarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena et al.ICML 2026
Builds on24
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
Related papers
- Understanding the Role of Training Data in Test-Time ScalingAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICLR 2026 · 5 citations
- T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference ScalingZhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang et al.ICML 2025
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLIan Wu, Yuxiao Qu, Amrith Setlur, Aviral KumarICML 2026 · 8 citations
- Atom of Thoughts for Markov LLM Test-Time ScalingFengwei Teng, Quan Shi, Zhaoyang Yu, Jiayi Zhang et al.NeurIPS 2025 · 73 citations
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length ConstraintYi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu et al.EMNLP 2025 · 1 citation
