Encoding Human Domain Knowledge to Warm Start Reinforcement Learning
Andrew Silva, Matthew C. Gombolay
Abstract
Deep reinforcement learning has been successful in a variety of tasks, such as game playing and robotic manipulation. However, attempting to learn tabula rasa disregards the logical structure of many domains as well as the wealth of readily available knowledge from domain experts that could help "warm start" the learning process. We present a novel reinforcement learning technique that allows for intelligent initialization of a neural network weights and architecture. Our approach permits the encoding domain knowledge directly into a neural decision tree, and improves upon that knowledge with policy gradient updates. We empirically validate our approach on two OpenAI Gym tasks and two modified StarCraft 2 tasks, showing that our novel architecture outperforms multilayer-perceptron and recurrent architectures. Our knowledge-based framework finds superior policies compared to imitation learning-based and prior knowledge-based approaches. Importantly, we demonstrate that our approach can be used by untrained humans to initially provide > 80% increase in expected reward relative to baselines prior to training (p < 0.001), which results in a > 60% increase in expected reward after policy optimization (p = 0.011).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- The Utility of Explainable AI in Ad Hoc Human-Machine TeamingRohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen et al.NeurIPS 2021 · 103 citations
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui et al.ICLR 2026 · 46 citations
- PAE: Reinforcement Learning from External Knowledge for Efficient ExplorationZhe Wu, Haofei Lu, Junliang Xing, You Wu et al.ICLR 2024 · 1 citation
- Mitigating Information Loss in Tree-Based Reinforcement Learning via Direct OptimizationSascha Marton, Tim Grams, Florian Vogt, Stefan Lüdtke et al.ICLR 2025
Builds on2
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsRohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. GombolayNeurIPS 2020 · 41 citations
Related papers
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 87 citations
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 6 citations
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du et al.NeurIPS 2020 · 48 citations
- Integrating Suboptimal Human Knowledge with Hierarchical Reinforcement Learning for Large-Scale Multiagent SystemsDingbang Liu, Shohei Kato, Wen Gu, Fenghui Ren et al.NeurIPS 2024 · 2 citations
- Creativity of AI: Automatic Symbolic Option Discovery for Facilitating Deep Reinforcement LearningMu Jin, Zhihao Ma, Kebing Jin, Hankz Hankui Zhuo et al.AAAI 2022 · 49 citations
