Encoding Human Domain Knowledge to Warm Start Reinforcement Learning
Andrew Silva, Matthew C. Gombolay
摘要
Deep reinforcement learning has been successful in a variety of tasks, such as game playing and robotic manipulation. However, attempting to learn tabula rasa disregards the logical structure of many domains as well as the wealth of readily available knowledge from domain experts that could help "warm start" the learning process. We present a novel reinforcement learning technique that allows for intelligent initialization of a neural network weights and architecture. Our approach permits the encoding domain knowledge directly into a neural decision tree, and improves upon that knowledge with policy gradient updates. We empirically validate our approach on two OpenAI Gym tasks and two modified StarCraft 2 tasks, showing that our novel architecture outperforms multilayer-perceptron and recurrent architectures. Our knowledge-based framework finds superior policies compared to imitation learning-based and prior knowledge-based approaches. Importantly, we demonstrate that our approach can be used by untrained humans to initially provide > 80% increase in expected reward relative to baselines prior to training (p < 0.001), which results in a > 60% increase in expected reward after policy optimization (p = 0.011).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The Utility of Explainable AI in Ad Hoc Human-Machine TeamingRohan R. Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen 等NeurIPS 2021 · 被引用 103 次
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui 等ICLR 2026 · 被引用 46 次
- PAE: Reinforcement Learning from External Knowledge for Efficient ExplorationZhe Wu, Haofei Lu, Junliang Xing, You Wu 等ICLR 2024 · 被引用 1 次
- Mitigating Information Loss in Tree-Based Reinforcement Learning via Direct OptimizationSascha Marton, Tim Grams, Florian Vogt, Stefan Lüdtke 等ICLR 2025
它引用的顶会 Paper2
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsRohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. GombolayNeurIPS 2020 · 被引用 41 次
相关 Paper
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 被引用 87 次
- Integrating Planning and Deep Reinforcement Learning via Automatic Induction of Task SubstructuresJung-Chun Liu, Chi-Hsien Chang, Shao-Hua Sun, Tian-Li YuICLR 2024 · 被引用 6 次
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du 等NeurIPS 2020 · 被引用 48 次
- Integrating Suboptimal Human Knowledge with Hierarchical Reinforcement Learning for Large-Scale Multiagent SystemsDingbang Liu, Shohei Kato, Wen Gu, Fenghui Ren 等NeurIPS 2024 · 被引用 2 次
- Creativity of AI: Automatic Symbolic Option Discovery for Facilitating Deep Reinforcement LearningMu Jin, Zhihao Ma, Kebing Jin, Hankz Hankui Zhuo 等AAAI 2022 · 被引用 49 次
