Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces
Haitong Ma, Ofir Nabati, Aviv Rosenberg, Bo Dai, Oran Lang, Craig Boutilier, Na Li, Shie Mannor, Lior Shani, Guy Tennenholtz
Abstract
Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a novel framework for training discrete diffusion models as highly effective policies in these complex settings. Our key innovation is an efficient online training process that ensures stable and effective policy improvement and . By leveraging policy mirror descent (PMD) to define an ideal, regularized target policy distribution, we frame the policy update as a distributional matching problem, training the expressive diffusion model to replicate this stable target. This decoupled approach stabilizes learning and significantly enhances training performance. Our method achieves state-of-the-art results and superior sample efficiency across a diverse set of challenging combinatorial benchmarks, including DNA sequence generation, RL with macro-actions, and multi-agent systems. Experiments demonstrate that our diffusion policies attain comparable or superior performance compared to other baselines. Crucially, our extensive empirical analysis reveals a key trade-off: FKL demonstrates superior sample efficiency and faster initial convergence, whereas RKL ensures stable training and higher asymptotic performance on challenging tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c93f53c6-fffb-4104-815d-ffd781650b70Cited by top-tier papers3
- Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial ActionsLingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti et al.ICML 2026 · 1 citation
- Trust-Region Diffusion Policies for Massively Parallel On-Policy RLHuy Le, Onur Celik, Denis Blessing, Tai Hoang et al.ICML 2026
- From Noise to Control: Parameterized Diffusion PoliciesRenhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris et al.ICML 2026
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
Related papers
- Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein DesignChenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang et al.ICLR 2025
- UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion ModelsJiaqi Wang, Haoge Deng, Ting Pan, Yang Liu et al.ICML 2026 · 3 citations
- Generative Online Reinforcement LearningChubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang et al.ICML 2026 · 1 citation
- Efficient Online Reinforcement Learning for Diffusion PolicyHaitong Ma, Tianyi Chen, Kai Wang, Na Li et al.ICML 2025
- Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion PoliciesZhuoran Li, Hai Zhong, Xun Wang, Qingxin Xia et al.ICML 2026 · 2 citations
