Neuro-algorithmic Policies Enable Fast Combinatorial Generalization
Marin Vlastelica P., Michal Rolínek, Georg Martius
Abstract
Although model-based and model-free approaches to learning the control of systems have achieved impressive results on standard benchmarks, generalization to task variations is still lacking. Recent results suggest that generalization for standard architectures improves only after obtaining exhaustive amounts of data. We give evidence that generalization capabilities are in many cases bottlenecked by the inability to generalize on the combinatorial aspects of the problem. Furthermore, we show that for a certain subclass of the MDP framework, this can be alleviated by neuro-algorithmic architectures. Many control problems require long-term planning that is hard to solve generically with neural networks alone. We introduce a neuro-algorithmic policy architecture consisting of a neural network and an embedded time-dependent shortest path solver. These policies can be trained end-to-end by blackbox differentiation. We show that this type of architecture generalizes well to unseen variations in the environment already after seeing a few examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3ddaff2-7d90-4dde-ae22-eb8ef2f0ff3eCited by top-tier papers6
- Explore to Generalize in Zero-Shot RLEv Zisselman, Itai Lavie, Daniel Soudry, Aviv TamarNeurIPS 2023 · 26 citations
- Optimize Planning Heuristics to Rank, not to Estimate Cost-to-GoalLeah Chrestien, Stefan Edelkamp, Antonín Komenda, Tomás PevnýNeurIPS 2023 · 17 citations
- Planning from Pixels in Environments with Combinatorially Hard Search SpacesMarco Bagatella, Miroslav Olsák, Michal Rolínek, Georg MartiusNeurIPS 2021 · 13 citations
- Causal Action Influence Aware Counterfactual Data AugmentationNúria Armengol Urpí, Marco Bagatella, Marin Vlastelica, Georg MartiusICML 2024 · 11 citations
- Backpropagation through Combinatorial Algorithms: Identity with Projection WorksSubham Sekhar Sahoo, Anselm Paulus, Marin Vlastelica, Vít Musil et al.ICLR 2023 · 11 citations
Builds on10
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius et al.ICLR 2020 · 341 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Smart Predict-and-Optimize for Hard Combinatorial Optimization ProblemsJayanta Mandi, Emir Demirovic, Peter J. Stuckey, Tias GunsAAAI 2020 · 184 citations
Related papers
- On the Expressivity of Neural Networks for Deep Reinforcement LearningKefan Dong, Yuping Luo, Tianhe Yu, Chelsea Finn et al.ICML 2020 · 33 citations
- On Computation and Reinforcement LearningRaj Ghugare, Michał Bortkiewicz, Alicja Ziarko, Benjamin EysenbachICML 2026
- Differentiable Weightless Controllers: Learning Logic Circuits for Continuous ControlFabian Kresse, Christoph LampertICML 2026 · 1 citation
- Neural Algorithmic Reasoning Without Intermediate SupervisionGleb Rodionov, Liudmila ProkhorenkovaNeurIPS 2023 · 20 citations
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 54 citations
