On the Modeling Capabilities of Large Language Models for Sequential Decision Making
Martin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan Mazoure
Abstract
Large pretrained models are showing increasingly better performance in reasoning and planning tasks across different modalities, opening the possibility to leverage them for complex sequential decision making problems. In this paper, we investigate the capabilities of Large Language Models (LLMs) for reinforcement learning (RL) across a diversity of interactive domains. We evaluate their ability to produce decision-making policies, either directly, by generating actions, or indirectly, by first generating reward models to train an agent with RL. Our results show that, even without task-specific fine-tuning, LLMs excel at reward modeling. In particular, crafting rewards through artificial intelligence (AI) feedback yields the most generally applicable approach and can enhance performance by improving credit assignment and exploration. Finally, in environments with unfamiliar dynamics, we explore how fine-tuning LLMs with synthetic data can significantly improve their reward modeling capabilities while mitigating catastrophic forgetting, further broadening their utility in sequential decision-making tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63f281a0-dadd-429e-a4d8-4daa2ccdefd8Cited by top-tier papers3
- Toward Efficient Exploration by Large Language Model AgentsDilip Arumugam, Thomas L. GriffithsICLR 2026 · 17 citations
- SkillGen: Learning Domain Skills for In-Context Sequential Decision MakingRuomeng Ding, Wei Cheng, Minglai Shao, Chen ZhaoAAAI 2026
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement LearningSilvia Sapora, R. Devon Hjelm, Omar Attia, Alexander Toshev et al.ICLR 2026
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- Efficient Sequential Decision Making with Large Language ModelsDingyang Chen, Qi Zhang, Yinglun ZhuEMNLP 2024 · 3 citations
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu et al.ICLR 2024 · 124 citations
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie et al.ICLR 2026 · 3 citations
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain GenerationJinbang Huang, Zhiyuan Li, Yuanzhao Hu, Zhanguang Zhang et al.ICML 2026 · 3 citations
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
