Discovering symbolic policies with deep reinforcement learning
Mikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago, Ruben Glatt, T. Nathan Mundhenk, Jacob F. Pettit, Daniel M. Faissol
摘要
Deep reinforcement learning (DRL) has proven successful for many difficult control problems by learning policies represented by neural networks. However, the complexity of neural network-based policies-involving thousands of composed nonlinear operators-can render them problematic to understand, trust, and deploy. In contrast, simple policies comprising short symbolic expressions can facilitate human understanding, while also being transparent and exhibiting predictable behavior. To this end, we propose deep symbolic policy, a novel approach to directly search the space of symbolic policies. We use an autoregressive recurrent neural network to generate control policies represented by tractable mathematical expressions, employing a risk-seeking policy gradient to maximize performance of the generated policies. To scale to environments with multidimensional action spaces, we propose an "anchoring" algorithm that distills pre-trained neural network-based policies into fully symbolic policies, one action dimension at a time. We also introduce two novel methods to improve exploration in DRL-based combinatorial optimization, building on ideas of entropy regularization and distribution initialization. Despite their dramatically reduced complexity, we demonstrate that discovered symbolic policies outperform seven state-of-the-art DRL algorithms in terms of average rank and average normalized episodic reward across eight benchmark environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- End-to-end Symbolic Regression with TransformersPierre-Alexandre Kamienny, Stéphane d'Ascoli, Guillaume Lample, François ChartonNeurIPS 2022 · 被引用 320 次
- A Unified Framework for Deep Symbolic RegressionMikel Landajuela, Chak Shing Lee, Jiachen Yang, Ruben Glatt 等NeurIPS 2022 · 被引用 160 次
- Transformer-based Planning for Symbolic RegressionParshin Shojaee, Kazem Meidani, Amir Barati Farimani, Chandan K. ReddyNeurIPS 2023 · 被引用 116 次
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- Symbolic Regression via Deep Reinforcement Learning Enhanced Genetic Programming SeedingT. Nathan Mundhenk, Mikel Landajuela, Ruben Glatt, Cláudio P. Santiago 等NeurIPS 2021 · 被引用 95 次
它引用的顶会 Paper1
相关 Paper
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi 等NeurIPS 2023 · 被引用 15 次
- Symbolic Distillation for Learned TCP Congestion ControlS. P. Sharan, Wenqing Zheng, Kuo-Feng Hsu, Jiarong Xing 等NeurIPS 2022 · 被引用 9 次
- Neurosymbolic Reinforcement Learning with Formally Verified ExplorationGreg Anderson, Abhinav Verma, Isil Dillig, Swarat ChaudhuriNeurIPS 2020 · 被引用 91 次
- A Neural-Guided Dynamic Symbolic Network for Exploring Mathematical Expressions from DataWenqiang Li, Weijun Li, Lina Yu, Min Wu 等ICML 2024 · 被引用 16 次
- Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery FrameworkYufei Kuang, Jie Wang, Haoyang Liu, Fangzhou Zhu 等ICLR 2024 · 被引用 15 次
