Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
Yunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rémi Munos
2025年份
16顶会引用
摘要
In this work, we investigate the merits of explicitly optimizing for inference time algorithmic performance during model training. We show how optimizing for inference time performance can improve overall model efficacy. We consider generic inference time objectives with k samples, with a focus on pass@k and majority voting as two main applications. With language model training on reasoning datasets, we showcase the performance trade-off enabled by training with such objectives. When training on code generation tasks, we show that the approach significantly improves pass@k objectives compared to the baseline method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Rewarding the Unlikely: Lifting GRPO Beyond Distribution SharpeningAndre Wang He, Daniel Fried, Sean WelleckEMNLP 2025 · 被引用 56 次
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
- Representation-Based Exploration for Language Models: From Test-Time to Post-TrainingJens Tuyls, Dylan J Foster, Akshay Krishnamurthy, Jordan T. AshICLR 2026 · 被引用 18 次
- Differential Smoothing Mitigates Sharpening and Improves LLM ReasoningJingchu Gai, Guanning Zeng, Huaqing ZHANG, Aditi RaghunathanICML 2026 · 被引用 13 次
- Polychromic Objectives for Reinforcement LearningJubayer Ibn Hamid, Ifdita Hasan Orney, Ellen Xu, Chelsea Finn 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- A Simple Model of Inference Scaling LawsNoam Itzhak LeviICML 2025
- Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical ReasoningFeng Chen, Allan Raventós, Nan Cheng, Surya Ganguli 等NeurIPS 2025 · 被引用 39 次
- OptScale: Probabilistic Optimality for Inference-time ScalingYoukang Wang, Jian Wang, Rubing Chen, Xiao-Yong WeiAAAI 2026 · 被引用 2 次
- Fixing the Broken Compass: Diagnosing and Improving Inference-Time Reward ModelingJiachun Li, Pengfei Cao, Zhuoran Jin, Yubo Chen 等ICLR 2026 · 被引用 3 次
