Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning
Renos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille, Wei Xia, Stefano Soatto
摘要
We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improved prediction with longer reasoning and larger models. Re-FORC enables: 1) early stopping of unpromising reasoning chains, reducing compute by up to 24% compared to fixed-budget cutoffs, while maintaining accuracy, 2) optimized model and thinking length selection that outperforms the largest model alone---reaching 1.7 percentage points higher peak accuracy while needing up to 12% less compute to match the largest model's accuracy, 3) adaptive test-time scaling, which increases accuracy by 8.8 percentage points (on average at maximum compute) over confidence-based baselines. Re-FORC allows dynamic reasoning with length control via cost-per-token thresholds while estimating computation time upfront.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
相关 Paper
- AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning ModelsJiacheng Wang, Tianle Chen, Pengyu Cheng, Xiaofeng Hou 等AAAI 2026
- Conformal Thinking: Risk Control for Reasoning on a Compute BudgetXi Wang, Anushri Suresh, Alvin Zhang, Rishi More 等ICML 2026 · 被引用 10 次
- Thought calibration: Efficient and confident test-time scalingMenghua Wu, Cai Zhou, Stephen Bates, Tommi S. JaakkolaEMNLP 2025 · 被引用 12 次
- SuCo: Sufficiency-guided Continuous Adaptive ReasoningJiahao Wang, Bingyu Liang, Chenhao Hu, Longhui Zhang 等ICML 2026
- Adaptive Thinking: Large Language Models Know When to Think in Latent SpacePingzhi Li, Bairu Hou, Yun Zhu, Yihao Feng 等ICLR 2026
