Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization Approach
Chanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing Zhang
Abstract
Large language models (LLMs) are increasingly deployed as "agents" for decision-making (DM) in interactive and dynamic environments. Yet, since they were not originally designed for DM, recent studies have shown that LLMs can struggle even in basic online DM problems, failing to exhibit desirable behaviors such as achieving a low regret value or an effective explorationexploitation tradeoff. To address this issue, we introduce Iterative Regret-Minimization Fine-Tuning (Iterative RMFT), a new post-training procedure that repeatedly distills low-regret decision trajectories back into the base model. At each iteration, the model rolls out several decision trajectories for a given online DM task. Iterative RMFT then selects the k-lowest regret trajectories, and trains the model on those trajectories via supervised fine-tuning. Unlike prior approaches that either (a) distill action sequences from known online DM algorithms, and/or (b) rely on manually crafted chain-of-thought outputs structured around known and fixed algorithms, our approach leverages the regret metric to automatically elicit the model's DM ability, and integrates the model's self-generated reasoning rationales. This reliance on model-generated reasoning avoids rigid output format engineering, and provides more flexible training signals in the natural language format. Empirical results show that Iterative RMFT can improve LLMs' DM performance across a spectrum of models-from Transformers with numerical input/output, to lightweight open-weight LLMs, and to the more advanced closed-weight LLM, GPT-4o mini. Notably, the flexibility that Iterative RMFT offers regarding output and reasoning formats allows the trained models to naturally exhibit generalization performance across tasks varying in time horizon, action space size, reward generation processes, and DM contexts/scenarios that are described by natural language. Finally, we provide a theoretical insight into how a single-layer attention Transformer model may lead to a no-regret learner under our training paradigm, in a simplified setting. Overall, we position our approach as an initial exploration, calling for more principled and novel post-training paradigms for LLMs when it comes to addressing DM tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 842df929-d731-4023-97c9-0b32bbbddb31Cited by top-tier papers1
Ask how each one uses itBuilds on27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
Related papers
- Do LLM Agents Have Regret? A Case Study in Online Learning and GamesChanwoo Park, Xiangyu Liu, Asuman E. Ozdaglar, Kaiqing ZhangICLR 2025
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie et al.ICLR 2026 · 3 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
- PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem SolvingMihir Parmar, Palash Goyal, Xin Liu, Yiwen Song et al.EMNLP 2025
