Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
Chanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing Zhang
摘要
Large language models (LLMs) are increasingly deployed as "agents" for decision-making (DM) in interactive and dynamic environments. Yet, since they were not originally designed for DM, recent studies have shown that LLMs can struggle even in basic online DM problems, failing to exhibit desirable behaviors such as achieving a low regret value or an effective explorationexploitation tradeoff. To address this issue, we introduce Iterative Regret-Minimization Fine-Tuning (Iterative RMFT), a new post-training procedure that repeatedly distills low-regret decision trajectories back into the base model. At each iteration, the model rolls out several decision trajectories for a given online DM task. Iterative RMFT then selects the k-lowest regret trajectories, and trains the model on those trajectories via supervised fine-tuning. Unlike prior approaches that either (a) distill action sequences from known online DM algorithms, and/or (b) rely on manually crafted chain-of-thought outputs structured around known and fixed algorithms, our approach leverages the regret metric to automatically elicit the model's DM ability, and integrates the model's self-generated reasoning rationales. This reliance on model-generated reasoning avoids rigid output format engineering, and provides more flexible training signals in the natural language format. Empirical results show that Iterative RMFT can improve LLMs' DM performance across a spectrum of models-from Transformers with numerical input/output, to lightweight open-weight LLMs, and to the more advanced closed-weight LLM, GPT-4o mini. Notably, the flexibility that Iterative RMFT offers regarding output and reasoning formats allows the trained models to naturally exhibit generalization performance across tasks varying in time horizon, action space size, reward generation processes, and DM contexts/scenarios that are described by natural language. Finally, we provide a theoretical insight into how a single-layer attention Transformer model may lead to a no-regret learner under our training paradigm, in a simplified setting. Overall, we position our approach as an initial exploration, calling for more principled and novel post-training paradigms for LLMs when it comes to addressing DM tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
相关 Paper
- Do LLM Agents Have Regret? A Case Study in Online Learning and GamesChanwoo Park, Xiangyu Liu, Asuman E. Ozdaglar, Kaiqing ZhangICLR 2025
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie 等ICLR 2026 · 被引用 3 次
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan 等NeurIPS 2024 · 被引用 214 次
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
- PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem SolvingMihir Parmar, Palash Goyal, Xin Liu, Yiwen Song 等EMNLP 2025
