e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Amrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer, Ian Wu, Virginia Smith, Max Simchowitz, Aviral Kumar
摘要
Test-time scaling offers a promising path to improve LLM reasoning by utilizing more compute at inference time; however, the true promise of this paradigm lies in extrapolation (i.e., improvement in performance on hard problems as LLMs keep "thinking" for longer, beyond the maximum token budget they were trained on). Surprisingly, we find that most existing reasoning models do not extrapolate well. We show that one way to enable extrapolation is by training the LLM to perform in-context exploration: training the LLM to effectively spend its test time budget by chaining operations (such as generation, verification, refinement, etc.), or testing multiple hypotheses before it commits to an answer. To enable in-context exploration, we identify three key ingredients as part of our recipe e3: (1) chaining skills that the base LLM has asymmetric competence in, e.g., chaining verification (easy) with generation (hard), as a way to implement in-context search; (2) leveraging "negative" gradients from incorrect traces to amplify exploration during RL, resulting in longer search traces that chains additional asymmetries; and (3) coupling task difficulty with training token budget during training via a specifically-designed curriculum to structure in-context exploration. Our recipe e3 produces the best known 1.7B model according to AIME'25 and HMMT'25 scores, and extrapolates to 2x the training token budget. Our e3-1.7B model not only attains high pass@1 scores, but also improves pass@k over the base model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- WebGym: Scaling Training Environments for Long-Horizon Visual Web Agents with Realistic TasksHao Bai, Alexey Taymanov, Tong Zhang, Aviral Kumar 等CVPR 2026
- Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsTejas Krishnan, Sumeet Motwani, Charles London, Suhaas Bhat 等ICML 2026
- OGPO: Sample Efficient Full-Finetuning of Generative Control PoliciesSarvesh Patil, Mitsuhiko Nakamoto, Manan Agarwal, Shashwat Saxena 等ICML 2026
它引用的顶会 Paper24
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
相关 Paper
- Understanding the Role of Training Data in Test-Time ScalingAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICLR 2026 · 被引用 5 次
- T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference ScalingZhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang 等ICML 2025
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLIan Wu, Yuxiao Qu, Amrith Setlur, Aviral KumarICML 2026 · 被引用 8 次
- Atom of Thoughts for Markov LLM Test-Time ScalingFengwei Teng, Quan Shi, Zhaoyang Yu, Jiayi Zhang 等NeurIPS 2025 · 被引用 73 次
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length ConstraintYi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu 等EMNLP 2025 · 被引用 1 次
