Black-Box Combinatorial Optimization with Order-Invariant Reinforcement Learning
Olivier Goudet, Quentin Suire, Adrien Goëffon, Frédéric Saubion, sylvain lamprier
Abstract
We introduce an order-invariant reinforcement learning framework for black-box combinatorial optimization. Classical estimation-of-distribution algorithms (EDAs) often rely on learning explicit variable dependency graphs, which can be costly and fail to capture complex interactions efficiently. In contrast, we parameterize a multivariate autoregressive generative model trained without a fixed variable ordering. By sampling random generation orders during training, a form of informationpreserving dropout, the model is encouraged to be invariant to variable order, promoting searchspace diversity, and shaping the model to focus on the most relevant variable dependencies, improving sample efficiency. We adapt Group Relative Policy Optimization (GRPO) to this setting, providing stable policy-gradient updates from scaleinvariant advantages. Across a wide range of benchmark algorithms and problem instances of varying sizes, our method frequently achieves the best performance and consistently avoids catastrophic failures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 239c6d9d-3a75-4976-a7b1-d969e1d4c136Builds on1
Related papers
- Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability ScalingAbhijeet Sinha, Sundari Elango, Dianbo LiuICML 2026 · 2 citations
- Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPOHongcheng Wang, Yinuo Huang, Sukai Wang, Guanghui Ren et al.ICML 2026 · 6 citations
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion ModelsTianren Ma, Mu Zhang, Yibing Wang, Qixiang YeICLR 2026 · 10 citations
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive GenerationShikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu et al.CVPR 2026 · 2 citations
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
