Mitigating Information Loss in Tree-Based Reinforcement Learning via Direct Optimization
Sascha Marton, Tim Grams, Florian Vogt, Stefan Lüdtke, Christian Bartelt, Heiner Stuckenschmidt
Abstract
Reinforcement learning (RL) has seen significant success across various domains, but its adoption is often limited by the black-box nature of neural network policies, making them difficult to interpret. In contrast, symbolic policies allow representing decision-making strategies in a compact and interpretable way. However, learning symbolic policies directly within on-policy methods remains challenging. In this paper, we introduce SYMPOL, a novel method for SYMbolic tree-based on-POLicy RL. SYMPOL employs a tree-based model integrated with a policy gradient method, enabling the agent to learn and adapt its actions while maintaining a high level of interpretability. We evaluate SYMPOL on a set of benchmark RL tasks, demonstrating its superiority over alternative tree-based RL approaches in terms of performance and interpretability. Unlike existing methods, it enables gradient-based, end-to-end learning of interpretable, axis-aligned decision trees within standard on-policy RL algorithms. Therefore, SYMPOL can become the foundation for a new class of interpretable RL based on decision trees. Our implementation is available under: https://github.com/s-marton/sympol * Equal Contribution Recently, the integration of symbolic methods into RL has gained significant attention. Symbolic RL does cover different approaches including program synthesis (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd16bbf6-5944-46ee-a632-da256a30948bCited by top-tier papers4
- Differentiable Decision Tree via "ReLU+Argmin" ReformulationQiangqiang Mao, Jiayang Ren, Yixiu Wang, Chenxuanyin Zou et al.NeurIPS 2025 · 2 citations
- Breiman meets Bellman: Non-Greedy Decision Trees with MDPsHector Kohler, Riad Akrour, Philippe PreuxKDD 2025
- Neural+Symbolic Approaches for Interpretable Actor-Critic Reinforcement LearningYue Yang, Fan Yang, Yu Bai, Hao WangICLR 2026
- Inducing, Detecting and Characterising Neural Modules: A Pipeline for Functional Interpretability in Reinforcement LearningAnna Soligo, Pietro Ferraro, David BoyleICML 2025
Builds on17
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
- Goal Misgeneralization in Deep Reinforcement LearningLauro Langosco di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau et al.ICML 2022 · 128 citations
- Discovering symbolic policies with deep reinforcement learningMikel Landajuela, Brenden K. Petersen, Sookyung Kim, Cláudio P. Santiago et al.ICML 2021 · 118 citations
- NODE-GAM: Neural Generalized Additive Model for Interpretable Deep LearningChun-Hao Chang, Rich Caruana, Anna GoldenbergICLR 2022 · 114 citations
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 104 citations
Related papers
- Efficient Symbolic Policy Learning with Differentiable Symbolic ExpressionJiaming Guo, Rui Zhang, Shaohui Peng, Qi Yi et al.NeurIPS 2023 · 15 citations
- BlendRL: A Framework for Merging Symbolic and Neural Policy LearningHikaru Shindo, Quentin Delfosse, Devendra Singh Dhami, Kristian KerstingICLR 2025
- Interpretable and Explainable Logical Policies via Neurally Guided Symbolic AbstractionQuentin Delfosse, Hikaru Shindo, Devendra Singh Dhami, Kristian KerstingNeurIPS 2023 · 64 citations
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual ExplanationsLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang et al.ICML 2024 · 18 citations
- Rethinking Branching on Exact Combinatorial Optimization Solver: The First Deep Symbolic Discovery FrameworkYufei Kuang, Jie Wang, Haoyang Liu, Fangzhou Zhu et al.ICLR 2024 · 15 citations
